Encoding device, decoding device, transmission device, and non-transitory storage medium

By employing a combination of block partition modes for dividing blocks into sub-blocks within the coding device, the signaling overhead is reduced, leading to improved compression efficiency in video and image coding.

JP2025085699AActive Publication Date: 2025-06-05PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2025039542
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-02-20
Filing Date
2025-03-12
Publication Date
2025-06-05
Estimated Expiration
2039-05-09

AI Technical Summary

Technical Problem

The increased signaling overhead due to deeper partition depths in block partitioning for video and image coding reduces video compression efficiency.

Method used

A coding device that uses a block partition determination unit to divide blocks into sub-blocks using a combination of block partition modes, including a first block partition mode defining a partition direction and number for the initial block division, and a second block division mode defining a division direction and number for further sub-block division, without signaling the number of divisions.

Benefits of technology

This approach improves compression efficiency by reducing the signaling overhead associated with deeper partition depths, thereby enhancing the effectiveness of video and image coding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025085699000001_ABST
    Figure 2025085699000001_ABST
Patent Text Reader

Abstract

To provide a transmission device capable of improving compression efficiency when encoding block division information.SOLUTION: An encoding device divides a block into multiple subblocks by using a block division mode set obtained by combining one or multiple block division modes that define a division type. When the number of the divisions in a first block division mode is 3, and a second block is a central block among the subblocks obtained after dividing the first block, and a division direction of a second block division mode is the same as a division direction of the first block division mode, a parameter for identifying the second block division mode includes a first flag indicating whether to divide the subblock horizontally or vertically but does not include a second flag indicating the number of divisions to divide the subblock.SELECTED DRAWING: Figure 33
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to methods and apparatus for encoding and decoding video and images using block partitioning. [Background technology]

[0002] In conventional image and video coding methods, images are typically divided into blocks, and coding and decoding processes are performed at the block level. Recent video standard developments allow coding and decoding processes with various block sizes in addition to the typical 8x8 or 16x16 sizes. A range of sizes from 4x4 to 256x256 can be used for image coding and decoding processes. [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] H.265(ISO / IEC 23008-2 HEVC(High Efficiency Video Coding)) Summary of the Invention [Problem to be solved by the invention]

[0004] To represent a range of sizes from 4x4 to 256x256, block partition information such as block partition mode (e.g., quadtree, binary tree, and ternary tree) and partition flag (e.g., split flag) is determined and signaled for a block. The signaling overhead increases as the partition depth increases. And the increased overhead reduces the video compression efficiency.

[0005] Therefore, an encoding device according to one aspect of the present disclosure provides a transmitting device etc. that can improve compression efficiency in encoding block division information. [Means for solving the problem]

[0006] A coding device according to one aspect of the present disclosure is a coding device for coding a picture, the coding device comprising a processor and a memory, the processor having, in operation, a block partition determination unit for dividing the acquired block into a plurality of sub-blocks using a block partition mode set that is a combination of one or more block partition modes that define a partition type, and a coding unit for coding the plurality of sub-blocks, the block partition mode set including a first block partition mode that defines a partition direction and a partition number for dividing a first block, and a coding unit for coding a second block that is one of the sub-blocks obtained after dividing the first block. and a second block division mode defining a division direction and number of divisions for dividing the first block into sub-blocks. When the number of divisions in the first block division mode is three, the second block is a central block among the sub-blocks obtained after dividing the first block, and the division direction of the second block division mode is the same as the division direction of the first block division mode, the parameters for identifying the second block division mode include a first flag indicating whether the sub-blocks are to be divided horizontally or vertically, and do not include a second flag indicating the number of divisions into which the sub-blocks are to be divided.

[0007] Furthermore, these comprehensive or specific aspects may be realized by a system, a method, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or may be realized by any combination of the system, the method, the integrated circuit, the computer program, and the recording medium. Effect of the Invention

[0008] According to the present disclosure, it is possible to improve compression efficiency in encoding block division information. [Brief description of the drawings]

[0009] [Figure 1] FIG. 1 is a block diagram showing a functional configuration of a coding device according to the first embodiment. [Diagram 2] FIG. 2 is a diagram showing an example of block division according to the first embodiment. [Diagram 3] FIG. 3 is a table showing the transform basis functions corresponding to each transform type. [Figure 4A] FIG. 4A is a diagram showing an example of the shape of a filter used in ALF. [Figure 4B] FIG. 4B is a diagram showing another example of the shape of the filter used in the ALF. [Figure 4C] FIG. 4C is a diagram showing another example of the shape of the filter used in the ALF. [Figure 5A] FIG. 5A is a diagram showing 67 intra prediction modes in intra prediction. [Figure 5B] FIG. 5B is a flowchart for explaining an outline of the predicted image correction process by the OBMC process. [Figure 5C] FIG. 5C is a conceptual diagram for explaining an overview of the predicted image correction process by the OBMC process. [Figure 5D] FIG. 5D is a diagram showing an example of FRUC. [Figure 6] FIG. 6 is a diagram for explaining pattern matching (bilateral matching) between two blocks along a motion trajectory. [Figure 7] FIG. 7 is a diagram for explaining pattern matching (template matching) between a template in a current picture and a block in a reference picture. [Figure 8] FIG. 8 is a diagram for explaining a model assuming uniform linear motion. [Figure 9A] FIG. 9A is a diagram for explaining derivation of a motion vector for each sub-block based on motion vectors of a plurality of adjacent blocks. [Figure 9B] FIG. 9B is a diagram for explaining an overview of the motion vector derivation process in the merge mode. [Figure 9C] FIG. 9C is a conceptual diagram for explaining an overview of the DMVR process. [Figure 9D]FIG. 9D is a diagram for explaining an outline of a predicted image generating method using luminance correction processing by LIC processing. [Figure 10] FIG. 10 is a block diagram showing a functional configuration of a decoding device according to the first embodiment. As shown in FIG. [Figure 11] FIG. 11 is a flowchart showing a video encoding process according to the second embodiment. [Figure 12] FIG. 12 is a flowchart showing a video decoding process according to the second embodiment. [Figure 13] FIG. 13 is a flowchart showing a video encoding process according to the third embodiment. [Figure 14] FIG. 14 is a flowchart showing a video decoding process according to the third embodiment. [Figure 15] FIG. 15 is a block diagram showing a structure of a video / image coding device according to the second or third embodiment. In FIG. [Figure 16] FIG. 16 is a block diagram showing a structure of a video / picture decoding device according to the second or third embodiment. In FIG. [Figure 17] FIG. 17 is a diagram showing examples of possible locations of the first parameter in a compressed video stream in the second or third embodiment. [Figure 18] FIG. 18 is a diagram showing examples of possible locations of the second parameter in the compressed video stream in the second or third embodiment. [Figure 19] FIG. 19 is a diagram showing an example of a second parameter following a first parameter in the second or third embodiment. [Figure 20] FIG. 20 is a diagram showing an example in which the second partition mode is not selected for division into blocks of 2NxN pixels as shown in step (2c) in the second embodiment. [Figure 21] FIG. 21 is a diagram showing an example in which the second partition mode is not selected for division into blocks of Nx2N pixels as shown in step (2c) in the second embodiment. [Figure 22]FIG. 22 is a diagram showing an example in which the second partition mode is not selected for division into blocks of NxN pixels as shown in step (2c) in the second embodiment. [Figure 23] FIG. 23 is a diagram showing an example in which the second partition mode is not selected for division into blocks of NxN pixels as shown in step (2c) in the second embodiment. [Figure 24] FIG. 24 is a diagram showing an example of dividing a block of 2N×N pixels using the partition mode selected when the second partition mode is not selected as shown in step (3) in the second embodiment. [Diagram 25] FIG. 25 is a diagram showing an example of dividing a block of Nx2N pixels using the partition mode selected when the second partition mode is not selected as shown in step (3) in the second embodiment. [Figure 26] FIG. 26 is a diagram showing an example of dividing a block of N×N pixels using the partition mode selected when the second partition mode is not selected as shown in step (3) in the second embodiment. [Figure 27] FIG. 27 is a diagram showing an example of dividing a block of N×N pixels using the partition mode selected when the second partition mode is not selected as shown in step (3) in the second embodiment. [Figure 28] 28A to 28H are diagrams showing examples of partition modes for dividing a block of NxN pixels in the second embodiment, where (a) to (h) are diagrams showing different partition modes. [Figure 29]29 is a diagram showing an example of partition types and partition directions for dividing a block of NxN pixels in embodiment 3. (1), (2), (3), and (4) are different partition types, (1a), (2a), (3a), and (4a) are partition modes with different partition types in the vertical partition direction, and (1b), (2b), (3b), and (4b) are partition modes with different partition types in the horizontal partition direction. [Diagram 30] FIG. 30 is a diagram showing an advantage of encoding the partition type before the partition direction compared to encoding the partition direction before the partition type in the third embodiment. [Figure 31A] FIG. 31A is a diagram showing an example of dividing a block into sub-blocks using a partition mode set with a smaller number of bins in partition mode encoding. [Figure 31B] FIG. 31B is a diagram showing an example of dividing a block into sub-blocks using a partition mode set with a smaller number of bins in partition mode encoding. [Figure 32A] FIG. 32A is a diagram showing an example of dividing a block into sub-blocks using a partition mode set that appears first in a predetermined order among a plurality of partition mode sets. [Figure 32B] FIG. 32B is a diagram showing an example of dividing a block into sub-blocks using a partition mode set that appears first in a predetermined order among a plurality of partition mode sets. [Figure 32C] FIG. 32C is a diagram showing an example of dividing a block into sub-blocks using a partition mode set that appears first in a predetermined order among a plurality of partition mode sets. [Diagram 33] FIG. 33 is a diagram showing the overall configuration of a content supply system that realizes a content distribution service. [Diagram 34] FIG. 34 is a diagram showing an example of a coding structure in scalable coding. [Diagram 35] FIG. 35 is a diagram showing an example of a coding structure in scalable coding. [Diagram 36] FIG. 36 is a diagram showing an example of a display screen of a web page. [Figure 37] FIG. 37 is a diagram showing an example of a display screen of a web page. [Figure 38] FIG. 38 is a diagram illustrating an example of a smartphone. [Figure 39] FIG. 39 is a block diagram showing an example of the configuration of a smartphone. [Diagram 40] FIG. 40 is a diagram showing an example of constraints in a partition mode in which a rectangular block is divided into three sub-blocks. [Diagram 41] FIG. 41 is a diagram showing an example of constraints in the partition mode in which a block is divided into two sub-blocks. [Diagram 42] FIG. 42 is a diagram showing an example of constraints in the partition mode in which a square block is divided into three sub-blocks. [Diagram 43] FIG. 43 is a diagram showing an example of constraints in a partition mode in which a rectangular block is divided into two sub-blocks. [Diagram 44] FIG. 44 is a diagram showing an example of constraints based on the division direction in a partition mode that divides a non-rectangular block into two sub-blocks. [Diagram 45] FIG. 45 is a diagram showing examples of valid partition directions for dividing a non-rectangular block into two sub-blocks. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] Hereinafter, the embodiment will be described in detail with reference to the drawings.

[0011] The embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, component arrangement and connection forms, steps, and order of steps shown in the following embodiments are merely examples and are not intended to limit the scope of the claims. Furthermore, among the components in the following embodiments, components that are not described in an independent claim showing a top concept are described as optional components.

[0012] (Embodiment 1) First, an overview of the first embodiment will be described as an example of an encoding device and a decoding device to which the processes and / or configurations described in each aspect of the present disclosure can be applied. However, the first embodiment is merely an example of an encoding device and a decoding device to which the processes and / or configurations described in each aspect of the present disclosure can be applied, and the processes and / or configurations described in each aspect of the present disclosure can also be implemented in encoding devices and decoding devices different from the first embodiment.

[0013] When applying the processing and / or configurations described in each aspect of the present disclosure to the first embodiment, for example, any of the following may be performed.

[0014] (1) For the encoding device or the decoding device of the first embodiment, among the multiple components constituting the encoding device or the decoding device, components corresponding to the components described in each aspect of the present disclosure are replaced with the components described in each aspect of the present disclosure. (2) In the encoding device or decoding device of the first embodiment, any modification such as addition, replacement, or deletion of functions or processes performed by some of the components constituting the encoding device or decoding device is made, and then the components corresponding to the components described in each aspect of the present disclosure are replaced with the components described in each aspect of the present disclosure. (3) Adding a process to the method implemented by the encoding device or decoding device of the first embodiment, and / or replacing or deleting some of the processes included in the method, and then replacing the process corresponding to the process described in each aspect of the present disclosure with the process described in each aspect of the present disclosure. (4) Some of the components constituting the encoding device or decoding device of the first embodiment may be implemented in combination with components described in each aspect of the present disclosure, components having some of the functions of the components described in each aspect of the present disclosure, or components performing some of the processing performed by the components described in each aspect of the present disclosure. (5) Implementing a component having some of the functions of some of the multiple components constituting the encoding device or decoding device of embodiment 1, or a component that performs some of the processing performed by some of the multiple components constituting the encoding device or decoding device of embodiment 1, in combination with a component described in each aspect of the present disclosure, a component having some of the functions of the component described in each aspect of the present disclosure, or a component that performs some of the processing performed by the component described in each aspect of the present disclosure. (6) In the method implemented by the encoding device or the decoding device of the first embodiment, among a plurality of processes included in the method, a process corresponding to a process described in each aspect of the present disclosure is replaced with a process described in each aspect of the present disclosure. (7) Some of the processes included in the method implemented by the encoding device or the decoding device of the first embodiment may be implemented in combination with the processes described in each aspect of the present disclosure.

[0015] It should be noted that the manner of implementing the processes and / or configurations described in each aspect of the present disclosure is not limited to the above examples. For example, the processes and / or configurations may be implemented in a device used for a purpose other than the video / image encoding device or video / image decoding device disclosed in the first embodiment, or the processes and / or configurations described in each aspect may be implemented independently. Furthermore, the processes and / or configurations described in different aspects may be implemented in combination.

[0016] [Outline of the encoding device] First, a description will be given of an overview of a coding device according to embodiment 1. Fig. 1 is a block diagram showing a functional configuration of a coding device 100 according to embodiment 1. The coding device 100 is a video / image coding device that codes a video / image on a block-by-block basis.

[0017] As shown in FIG. 1, the encoding device 100 is a device that encodes an image on a block-by-block basis, and includes a division unit 102, a subtraction unit 104, a transformation unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse transformation unit 114, an addition unit 116, a block memory 118, a loop filter unit 120, a frame memory 122, an intra prediction unit 124, an inter prediction unit 126, and a prediction control unit 128.

[0018] The encoding device 100 is realized by, for example, a general-purpose processor and a memory. In this case, when the software program stored in the memory is executed by the processor, the processor functions as the division unit 102, the subtraction unit 104, the transformation unit 106, the quantization unit 108, the entropy coding unit 110, the inverse quantization unit 112, the inverse transformation unit 114, the addition unit 116, the loop filter unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128. The encoding device 100 may also be realized as one or more dedicated electronic circuits corresponding to the division unit 102, the subtraction unit 104, the transformation unit 106, the quantization unit 108, the entropy coding unit 110, the inverse quantization unit 112, the inverse transformation unit 114, the addition unit 116, the loop filter unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128.

[0019] Each component included in the encoding device 100 will be described below.

[0020] [Divided part] The division unit 102 divides each picture included in the input video into a plurality of blocks, and outputs each block to the subtraction unit 104. For example, the division unit 102 first divides a picture into blocks of a fixed size (e.g., 128x128). These fixed-size blocks may be called coding tree units (CTUs). The division unit 102 then divides each of the fixed-size blocks into blocks of a variable size (e.g., 64x64 or less) based on recursive quadtree and / or binary tree block division. These variable-size blocks may be called coding units (CUs), prediction units (PUs), or transform units (TUs). Note that in this embodiment, it is not necessary to distinguish between CUs, PUs, and TUs, and some or all of the blocks in a picture may be the processing units of CUs, PUs, and TUs.

[0021] Fig. 2 is a diagram showing an example of block division according to embodiment 1. In Fig. 2, solid lines represent block boundaries based on quadtree block division, and dashed lines represent block boundaries based on binary tree block division.

[0022] Here, the block 10 is a square block of 128x128 pixels (128x128 block). This 128x128 block 10 is first divided into four square 64x64 blocks (quadtree block division).

[0023] The top-left 64x64 block is further divided vertically into two rectangular 32x64 blocks, and the left 32x64 block is further divided vertically into two rectangular 16x64 blocks (binary tree block division). As a result, the top-left 64x64 block is divided into two 16x64 blocks 11 and 12 and a 32x64 block 13.

[0024] The top right 64x64 block is divided horizontally into two rectangular 64x32 blocks 14, 15 (binary tree block division).

[0025] The bottom left 64x64 block is divided into four square 32x32 blocks (quadtree block division). Of the four 32x32 blocks, the top left and bottom right blocks are further divided. The top left 32x32 block is divided vertically into two rectangular 16x32 blocks, and the right 16x32 block is further divided horizontally into two 16x16 blocks (binary tree block division). The bottom right 32x32 block is divided horizontally into two 32x16 blocks (binary tree block division). As a result, the bottom left 64x64 block is divided into a 16x32 block 16, two 16x16 blocks 17, 18, two 32x32 blocks 19, 20, and two 32x16 blocks 21, 22.

[0026] The bottom right 64x64 block 23 is not split.

[0027] 2, the block 10 is divided into 13 variable-sized blocks 11 to 23 based on recursive quad-tree and binary tree block division. Such division is sometimes called QTBT (quad-tree plus binary tree) division.

[0028] In Fig. 2, one block is divided into four or two blocks (quadtree or binary tree block division), but the division is not limited to this. For example, one block may be divided into three blocks (ternary tree block division). Division including such ternary tree block division is sometimes called MBT (multi type tree) division.

[0029] [Subtraction section] The subtraction unit 104 subtracts a prediction signal (prediction sample) from an original signal (original sample) for each block divided by the division unit 102. That is, the subtraction unit 104 calculates a prediction error (also called a residual error) of a block to be coded (hereinafter, referred to as a current block). Then, the subtraction unit 104 outputs the calculated prediction error to the conversion unit 106.

[0030] The original signal is an input signal to the encoding device 100, and is a signal representing an image of each picture constituting a moving image (for example, a luminance (luma) signal and two color difference (chroma) signals). Hereinafter, the signal representing the image may also be referred to as a sample.

[0031] [Conversion section] The transform unit 106 transforms the prediction error in the spatial domain into transform coefficients in the frequency domain, and outputs the transform coefficients to the quantization unit 108. Specifically, the transform unit 106 performs, for example, a predetermined discrete cosine transform (DCT) or discrete sine transform (DST) on the prediction error in the spatial domain.

[0032] The transform unit 106 may adaptively select a transform type from among a plurality of transform types, and transform the prediction errors into transform coefficients using a transform basis function corresponding to the selected transform type. Such a transform may be called an explicit multiple core transform (EMT) or an adaptive multiple transform (AMT).

[0033] The multiple transform types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. Figure 3 is a table showing the transform basis functions corresponding to each transform type. In Figure 3, N indicates the number of input pixels. The selection of the transform type from among the multiple transform types may depend on, for example, the type of prediction (intra prediction and inter prediction) or the intra prediction mode.

[0034] Such information indicating whether EMT or AMT is applied (e.g., called an AMT flag) and information indicating the selected transformation type are signaled at the CU level. Note that the signaling of such information does not need to be limited to the CU level, and may be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0035] Furthermore, the transform unit 106 may retransform the transform coefficients (transformation results). Such retransformation may be called AST (adaptive secondary transform) or NSST (non-separable secondary transform). For example, the transform unit 106 performs retransformation for each subblock (e.g., 4x4 subblock) included in a block of transform coefficients corresponding to intra-prediction errors. Information indicating whether or not to apply NSST and information regarding a transform matrix used in NSST are signaled at a CU level. Note that signaling of these pieces of information does not need to be limited to the CU level, and may be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0036] Here, a separable transformation is a method in which the transformation is performed multiple times by separating the input into directions equal to the number of dimensions, and a non-separable transformation is a method in which when the input is multidimensional, two or more dimensions are treated as one dimension and the transformation is performed together.

[0037] For example, one example of a non-separable transformation is when the input is a 4x4 block, it is treated as a single array with 16 elements, and the transformation process is performed on that array using a 16x16 transformation matrix.

[0038] Another example of a non-separable transformation is the Hypercube Givens Transform, which treats a 4x4 input block as a single array with 16 elements and then performs Givens rotations on that array multiple times.

[0039] [Quantization section] The quantization unit 108 quantizes the transform coefficients output from the transform unit 106. Specifically, the quantization unit 108 scans the transform coefficients of the current block in a predetermined scanning order, and quantizes the transform coefficients based on a quantization parameter (QP) corresponding to the scanned transform coefficients. Then, the quantization unit 108 outputs the quantized transform coefficients of the current block (hereinafter, referred to as quantized coefficients) to the entropy coding unit 110 and the inverse quantization unit 112.

[0040] The predetermined order is an order for quantization / dequantization of the transform coefficients. For example, the predetermined scanning order is defined as ascending (low to high) or descending (high to low) frequency order.

[0041] The quantization parameter is a parameter that defines the quantization step (quantization width). For example, if the value of the quantization parameter increases, the quantization step also increases. In other words, if the value of the quantization parameter increases, the quantization error increases.

[0042] [Entropy coding part] The entropy coding unit 110 generates a coded signal (coded bit stream) by variable-length coding the quantized coefficients input from the quantization unit 108. Specifically, the entropy coding unit 110, for example, binarizes the quantized coefficients and arithmetically codes the binary signal.

[0043] [Dequantization section] The inverse quantization unit 112 inverse quantizes the quantized coefficients input from the quantization unit 108. Specifically, the inverse quantization unit 112 inverse quantizes the quantized coefficients of the current block in a predetermined scanning order. Then, the inverse quantization unit 112 outputs the inverse quantized transform coefficients of the current block to the inverse transform unit 114.

[0044] [Inverse conversion section] The inverse transform unit 114 restores the prediction error by inverse transforming the transform coefficients that are input from the inverse quantization unit 112. Specifically, the inverse transform unit 114 restores the prediction error of the current block by performing an inverse transform on the transform coefficients that corresponds to the transform performed by the transform unit 106. Then, the inverse transform unit 114 outputs the restored prediction error to the adder unit 116.

[0045] Note that the restored prediction error does not match the prediction error calculated by the subtraction unit 104 because information has been lost due to quantization. That is, the restored prediction error includes a quantization error.

[0046] [Addition section] The adder 116 reconstructs the current block by adding the prediction error input from the inverse transformer 114 and the prediction sample input from the prediction control unit 128. The adder 116 then outputs the reconstructed block to the block memory 118 and the loop filter unit 120. The reconstructed block is sometimes called a local decoded block.

[0047] [Block memory] The block memory 118 is a storage unit for storing blocks that are referenced in intra prediction and are in a picture to be coded (hereinafter, referred to as a current picture). Specifically, the block memory 118 stores the reconstructed blocks output from the adder 116.

[0048] [Loop filter section] The loop filter unit 120 applies a loop filter to the block reconstructed by the adder unit 116, and outputs the filtered reconstructed block to the frame memory 122. The loop filter is a filter (in-loop filter) used in the encoding loop, and includes, for example, a deblocking filter (DF), a sample adaptive offset (SAO), and an adaptive loop filter (ALF).

[0049] In ALF, a least squared error filter is applied to remove coding artifacts. For example, for each 2x2 sub-block in the current block, one filter is selected from among multiple filters based on local gradient direction and activity.

[0050] Specifically, first, sub-blocks (e.g., 2x2 sub-blocks) are classified into a plurality of classes (e.g., 15 or 25 classes). The classification of the sub-blocks is performed based on the gradient direction and activity. For example, a classification value C (e.g., C=5D+A) is calculated using a gradient direction value D (e.g., 0 to 2 or 0 to 4) and a gradient activity value A (e.g., 0 to 4). Then, based on the classification value C, the sub-blocks are classified into a plurality of classes (e.g., 15 or 25 classes).

[0051] The gradient direction value D is derived, for example, by comparing gradients in multiple directions (e.g., horizontal, vertical, and two diagonal directions), and the gradient activity value A is derived, for example, by adding gradients in multiple directions and quantizing the sum.

[0052] Based on the result of such classification, a filter for the sub-block is determined from among a plurality of filters.

[0053] The shape of the filter used in the ALF is, for example, a circularly symmetric shape. FIGS. 4A to 4C are diagrams showing a number of examples of the shape of the filter used in the ALF. FIG. 4A shows a 5×5 diamond-shaped filter, FIG. 4B shows a 7×7 diamond-shaped filter, and FIG. 4C shows a 9×9 diamond-shaped filter. Information indicating the shape of the filter is signaled at the picture level. Note that the signaling of the information indicating the shape of the filter does not need to be limited to the picture level, and may be at other levels (for example, the sequence level, slice level, tile level, CTU level, or CU level).

[0054] The on / off of ALF is determined, for example, at the picture level or the CU level. For example, whether or not to apply ALF is determined for luminance at the CU level, and whether or not to apply ALF is determined for chrominance at the picture level. Information indicating whether or not to apply ALF is signaled at the picture level or the CU level. Note that the signaling of information indicating whether or not to apply ALF is not limited to the picture level or the CU level, and may be at another level (for example, the sequence level, the slice level, the tile level, or the CTU level).

[0055] The coefficient sets of multiple selectable filters (e.g., up to 15 or 25 filters) are signaled at the picture level. Note that the signaling of the coefficient sets does not need to be limited to the picture level, but may be at other levels (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).

[0056] [Frame memory] The frame memory 122 is a storage unit for storing reference pictures used in inter prediction, and may be called a frame buffer. Specifically, the frame memory 122 stores the reconstructed block filtered by the loop filter unit 120.

[0057] [Intra prediction section] The intra prediction unit 124 generates a prediction signal (intra prediction signal) by performing intra prediction (also called intra-screen prediction) of the current block with reference to a block in the current picture stored in the block memory 118. Specifically, the intra prediction unit 124 generates an intra prediction signal by performing intra prediction with reference to samples (e.g., luminance values, chrominance values) of blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control unit 128.

[0058] For example, the intra prediction unit 124 performs intra prediction using one of a plurality of predefined intra prediction modes. The plurality of intra prediction modes includes one or more non-directional prediction modes and a plurality of directional prediction modes.

[0059] The one or more non-directional prediction modes include, for example, a planar prediction mode and a DC prediction mode defined in the H.265 / High-Efficiency Video Coding (HEVC) standard (Non-Patent Document 1).

[0060] The multiple directional prediction modes include, for example, 33 prediction modes defined in the H.265 / HEVC standard. The multiple directional prediction modes may include 32 prediction modes in addition to the 33 directions (a total of 65 directional prediction modes). FIG. 5A is a diagram showing 67 intra prediction modes (2 non-directional prediction modes and 65 directional prediction modes) in intra prediction. The solid arrows represent the 33 directions defined in the H.265 / HEVC standard, and the dashed arrows represent the additional 32 directions.

[0061] In addition, in the intra prediction of the chrominance block, the luminance block may be referenced. That is, the chrominance component of the current block may be predicted based on the luminance component of the current block. Such intra prediction may be called CCLM (cross-component linear model) prediction. An intra prediction mode of the chrominance block that refers to such a luminance block (for example, called a CCLM mode) may be added as one of the intra prediction modes of the chrominance block.

[0062] The intra prediction unit 124 may correct pixel values ​​after intra prediction based on the gradient of reference pixels in the horizontal / vertical directions. Intra prediction with such correction may be called position dependent intra prediction combination (PDPC). Information indicating whether or not PDPC is applied (e.g., called a PDPC flag) is signaled, for example, at a CU level. Note that the signaling of this information does not need to be limited to the CU level, and may be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0063] [Inter prediction section] The inter prediction unit 126 generates a prediction signal (inter prediction signal) by performing inter prediction (also called inter prediction) of the current block with reference to a reference picture stored in the frame memory 122 and different from the current picture. The inter prediction is performed in units of the current block or a sub-block (e.g., 4x4 block) in the current block. For example, the inter prediction unit 126 performs motion estimation in the reference picture for the current block or the sub-block. Then, the inter prediction unit 126 generates an inter prediction signal of the current block or the sub-block by performing motion compensation using motion information (e.g., a motion vector) obtained by the motion estimation. Then, the inter prediction unit 126 outputs the generated inter prediction signal to the prediction control unit 128.

[0064] The motion information used for the motion compensation is signaled. For the signaling of the motion vector, a motion vector predictor may be used, i.e. the difference between the motion vector and the motion vector predictor may be signaled.

[0065] In addition, the inter prediction signal may be generated using not only the motion information of the current block obtained by motion search, but also the motion information of the adjacent block. Specifically, the inter prediction signal may be generated for each sub-block in the current block by performing weighted addition of the prediction signal based on the motion information obtained by motion search and the prediction signal based on the motion information of the adjacent block. Such inter prediction (motion compensation) may be called OBMC (overlapped block motion compensation).

[0066] In such an OBMC mode, information indicating the size of a sub-block for OBMC (e.g., called OBMC block size) is signaled at the sequence level. Also, information indicating whether or not to apply the OBMC mode (e.g., called OBMC flag) is signaled at the CU level. Note that the signaling level of these pieces of information does not need to be limited to the sequence level and CU level, and may be other levels (e.g., picture level, slice level, tile level, CTU level, or sub-block level).

[0067] The OBMC mode will now be described in more detail. Figures 5B and 5C are a flowchart and a conceptual diagram for explaining an overview of the predicted image correction process in the OBMC process.

[0068] First, a predicted image (Pred) is obtained by normal motion compensation using a motion vector (MV) assigned to a block to be coded.

[0069] Next, the motion vector (MV_L) of the already-encoded left adjacent block is applied to the block to be encoded to obtain a predicted image (Pred_L), and the predicted image is weighted and superimposed with Pred_L to perform a first correction of the predicted image.

[0070] Similarly, the motion vector (MV_U) of the already-encoded adjacent block above is applied to the block to be encoded to obtain a predicted image (Pred_U), and the predicted image that has been corrected the first time is weighted and overlaid with Pred_U to perform a second correction of the predicted image, which is used as the final predicted image.

[0071] Although a two-stage correction method using the left adjacent block and the upper adjacent block has been described here, it is also possible to configure a method in which more than two stages of correction are performed using the right adjacent block or the lower adjacent block.

[0072] The area in which overlapping is performed does not have to be the entire pixel area of ​​the block, but may be only a part of the area near the block boundary.

[0073] Note that although the predicted image correction process from one reference picture has been described here, the process is similar when correcting a predicted image from multiple reference pictures; after obtaining corrected predicted images from each reference picture, the obtained predicted images are further overlaid to obtain the final predicted image.

[0074] The block to be processed may be a prediction block unit, or a sub-block unit obtained by further dividing the prediction block.

[0075] As a method of determining whether or not to apply OBMC processing, for example, there is a method of using obmc_flag, which is a signal indicating whether or not to apply OBMC processing. As a specific example, in an encoding device, it is determined whether or not the encoding target block belongs to an area with complex motion, and if it belongs to an area with complex motion, a value of 1 is set as obmc_flag and encoding is performed by applying OBMC processing, and if it does not belong to an area with complex motion, a value of 0 is set as obmc_flag and encoding is performed without applying OBMC processing. On the other hand, in a decoding device, by decoding obmc_flag described in a stream, whether or not to apply OBMC processing is switched according to the value, and decoding is performed.

[0076] In addition, the motion information may be derived on the decoding device side without being signaled. For example, a merge mode defined in the H.265 / HEVC standard may be used. Also, for example, the motion information may be derived by performing motion estimation on the decoding device side. In this case, the motion estimation is performed without using pixel values ​​of the current block.

[0077] Here, a mode in which motion estimation is performed on the decoding device side will be described. This mode in which motion estimation is performed on the decoding device side is sometimes called a pattern matched motion vector derivation (PMMVD) mode or a frame rate up-conversion (FRUC) mode.

[0078] An example of the FRUC process is shown in FIG. 5D. First, a list of multiple candidates (which may be the same as the merge list) each having a predicted motion vector is generated by referring to the motion vectors of encoded blocks spatially or temporally adjacent to the current block. Next, a best candidate MV is selected from multiple candidate MVs registered in the candidate list. For example, an evaluation value of each candidate included in the candidate list is calculated, and one candidate is selected based on the evaluation value.

[0079] Then, based on the motion vector of the selected candidate, a motion vector for the current block is derived. Specifically, for example, the motion vector of the selected candidate (best candidate MV) is derived as it is as the motion vector for the current block. Also, for example, the motion vector for the current block may be derived by performing pattern matching in the surrounding area of ​​the position in the reference picture corresponding to the motion vector of the selected candidate. That is, a search is performed in the surrounding area of ​​the best candidate MV in the same manner, and if there is an MV with a better evaluation value, the best candidate MV may be updated to the MV, and the MV may be set as the final MV of the current block. It is also possible to configure the system without performing this process.

[0080] The same processing may be performed when processing is performed in sub-block units.

[0081] The evaluation value is calculated by finding a difference value of the reconstructed image by pattern matching between an area in a reference picture corresponding to the motion vector and a predetermined area. The evaluation value may be calculated using information other than the difference value.

[0082] As the pattern matching, a first pattern matching or a second pattern matching is used. The first pattern matching and the second pattern matching are sometimes called bilateral matching and template matching, respectively.

[0083] In the first pattern matching, pattern matching is performed between two blocks in two different reference pictures that are along the motion trajectory of the current block. Therefore, in the first pattern matching, an area in another reference picture along the motion trajectory of the current block is used as a predetermined area for calculating the evaluation value of the above-mentioned candidate.

[0084] FIG. 6 is a diagram for explaining an example of pattern matching (bilateral matching) between two blocks along a motion trajectory. As shown in FIG. 6, in the first pattern matching, two motion vectors (MV0, MV1) are derived by searching for a pair of two blocks that are along the motion trajectory of a current block (Cur block) and are in two different reference pictures (Ref0, Ref1) that best match each other. Specifically, for the current block, a difference is derived between a reconstructed image at a designated position in a first coded reference picture (Ref0) designated by a candidate MV and a reconstructed image at a designated position in a second coded reference picture (Ref1) designated by a symmetric MV obtained by scaling the candidate MV by a display time interval, and an evaluation value is calculated using the obtained difference value. It is preferable to select the candidate MV with the best evaluation value among a plurality of candidate MVs as the final MV.

[0085] Under the assumption of continuous motion trajectories, the motion vectors (MV0, MV1) pointing to two reference blocks are proportional to the temporal distances (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, if the current picture is located between two reference pictures in time and the temporal distances from the current picture to the two reference pictures are equal, the first pattern matching derives bidirectional motion vectors that are mirror-symmetric.

[0086] In the second pattern matching, pattern matching is performed between a template in the current picture (a block adjacent to the current block in the current picture (e.g., an upper and / or left adjacent block)) and a block in the reference picture. Therefore, in the second pattern matching, a block adjacent to the current block in the current picture is used as a predetermined area for calculating the evaluation value of the above-mentioned candidate.

[0087] Fig. 7 is a diagram for explaining an example of pattern matching (template matching) between a template in a current picture and a block in a reference picture. As shown in Fig. 7, in the second pattern matching, a motion vector of a current block is derived by searching in a reference picture (Ref0) for a block that best matches a block adjacent to a current block (Cur block) in a current picture (Cur Pic). Specifically, for the current block, a difference is derived between a reconstructed image of both or either of the left adjacent and / or upper adjacent coded areas and a reconstructed image at the same position in a coded reference picture (Ref0) specified by a candidate MV, an evaluation value is calculated using the obtained difference value, and a candidate MV with the best evaluation value among a plurality of candidate MVs is selected as a best candidate MV.

[0088] Information indicating whether such a FRUC mode is applied (e.g., when the FRUC flag is true) is signaled at the CU level. Also, when the FRUC mode is applied (e.g., when the FRUC flag is true), information indicating a pattern matching method (first pattern matching or second pattern matching) (e.g., when the FRUC mode flag is true) is signaled at the CU level. Note that the signaling of such information does not need to be limited to the CU level, and may be at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or subblock level).

[0089] Here, a mode for deriving a motion vector based on a model assuming uniform linear motion will be described. This mode is based on BIO (bi-directional optical This is sometimes called "flow" mode.

[0090] Fig. 8 is a diagram for explaining a model assuming uniform linear motion. In Fig. 8, (vx, vy) indicates a velocity vector, and τ0 and τ1 indicate the temporal distance between the current picture (Cur Pic) and two reference pictures (Ref0, Ref1), respectively. (MVx0, MVy0) indicates a motion vector corresponding to the reference picture Ref0, and (MVx1, MVy1) indicates a motion vector corresponding to the reference picture Ref1.

[0091] In this case, under the assumption of uniform linear motion of the velocity vector (vx, vy), (MVx0, MVy0) and (MVx1, MVy1) are expressed as (vxτ0, vyτ0) and (-vxτ1, -vyτ1), respectively, and the following optical flow equation (1) holds.

[0092]

number

[0093] Here, I(k) denotes the luminance value of reference image k (k=0,1) after motion compensation. This optical flow equation indicates that the sum of (i) the time derivative of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Based on a combination of this optical flow equation and Hermite interpolation, block-wise motion vectors obtained from a merge list or the like are corrected pixel by pixel.

[0094] Note that the decoding device may derive a motion vector using a method other than the method based on a model assuming uniform linear motion. For example, a motion vector may be derived for each sub-block based on the motion vectors of multiple adjacent blocks.

[0095] Here, a mode in which a motion vector is derived for each sub-block based on the motion vectors of a plurality of adjacent blocks will be described. This mode is sometimes called an affine motion compensation prediction mode.

[0096] Fig. 9A is a diagram for explaining derivation of a motion vector for each subblock based on the motion vectors of multiple adjacent blocks. In Fig. 9A, the current block includes 16 4x4 subblocks. Here, the motion vector v0 of the upper left corner control point of the current block is derived based on the motion vectors of the adjacent blocks, and the motion vector v1 of the upper right corner control point of the current block is derived based on the motion vectors of the adjacent subblocks. Then, using the two motion vectors v0 and v1, the motion vector (vx, vy) of each subblock in the current block is derived by the following formula (2).

[0097]

number

[0098] Here, x and y respectively indicate the horizontal and vertical positions of the sub-block, and w indicates a predetermined weighting factor.

[0099] Such affine motion compensation prediction mode may include several modes with different methods of deriving the motion vectors of the upper left and upper right corner control points. Information indicating such affine motion compensation prediction mode (e.g., called affine flag) is signaled at the CU level. Note that the signaling of the information indicating this affine motion compensation prediction mode does not need to be limited to the CU level, and may be at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or subblock level).

[0100] [Predictive control unit] The prediction control unit 128 selects either the intra-prediction signal or the inter-prediction signal, and outputs the selected signal to the subtraction unit 104 and the addition unit 116 as a prediction signal.

[0101] Here, an example of deriving a motion vector for a picture to be coded in the merge mode will be described. Fig. 9B is a diagram for explaining an overview of a motion vector derivation process in the merge mode.

[0102] First, a prediction MV list is generated in which prediction MV candidates are registered. The prediction MV candidates include spatially adjacent prediction MVs, which are MVs held by multiple coded blocks located spatially around the block to be coded, temporally adjacent prediction MVs, which are MVs held by nearby blocks projected from the position of the block to be coded in the coded reference picture, joint prediction MVs, which are MVs generated by combining the MV values ​​of spatially adjacent prediction MVs and temporally adjacent prediction MVs, and zero prediction MVs, which are MVs with a value of zero.

[0103] Next, one prediction MV is selected from the multiple prediction MVs registered in the prediction MV list, and is determined as the MV for the block to be coded.

[0104] Furthermore, the variable length coding unit writes merge_idx, which is a signal indicating which predicted MV has been selected, into the stream and codes it.

[0105] Note that the predicted MVs registered in the predicted MV list described in Figure 9B are just an example, and the number may be different from the number shown in the figure, the configuration may not include some of the types of predicted MVs shown in the figure, or the configuration may include additional predicted MVs other than the types of predicted MVs shown in the figure.

[0106] Note that the final MV may be determined by performing DMVR processing, which will be described later, using the MV of the block to be coded derived in the merge mode.

[0107] Here, an example of determining the MV using the DMVR process will be described.

[0108] FIG. 9C is a conceptual diagram for explaining an overview of the DMVR process.

[0109] First, the optimal MVP set for the block to be processed is set as a candidate MV, and reference pixels are obtained from the first reference picture, which is a processed picture in the L0 direction, and the second reference picture, which is a processed picture in the L1 direction, according to the candidate MV, and a template is generated by taking the average of each reference pixel.

[0110] Next, the template is used to search the surrounding areas of the candidate MVs of the first and second reference pictures, and the MV with the smallest cost is determined as the final MV. The cost value is calculated using the difference value between each pixel value of the template and each pixel value of the search area, the MV value, etc.

[0111] The outline of the processing described here is basically the same for the encoding device and the decoding device.

[0112] Note that other processing may be used instead of the processing described here, as long as it is processing that can search the vicinity of the candidate MV and derive the final MV.

[0113] Here, a mode in which a predicted image is generated using LIC processing will be described.

[0114] FIG. 9D is a diagram for explaining an outline of a predicted image generating method using luminance correction processing by LIC processing.

[0115] First, a MV for obtaining a reference image corresponding to a block to be coded from a reference picture that is a coded picture is derived.

[0116] Next, for the block to be coded, the luminance pixel values ​​of the coded surrounding reference areas adjacent to the left and above and the luminance pixel values ​​at the equivalent positions in the reference picture specified by the MV are used to extract information indicating how the luminance values ​​have changed between the reference picture and the picture to be coded, and a luminance correction parameter is calculated.

[0117] A luminance correction process is performed on a reference image in a reference picture specified by the MV using the luminance correction parameter, thereby generating a predicted image for the block to be coded.

[0118] It should be noted that the shape of the peripheral reference region in FIG. 9D is just an example, and other shapes may be used.

[0119] Although the process of generating a predicted image from one reference picture has been described here, the process is similar when generating a predicted image from multiple reference pictures, and a luminance correction process is performed in a similar manner on the reference images obtained from each reference picture before generating a predicted image.

[0120] As a method of determining whether or not to apply LIC processing, for example, there is a method of using lic_flag, which is a signal indicating whether or not to apply LIC processing. As a specific example, in an encoding device, it is determined whether or not the encoding target block belongs to an area where a luminance change occurs, and if it belongs to an area where a luminance change occurs, a value of 1 is set as lic_flag and encoding is performed by applying LIC processing, and if it does not belong to an area where a luminance change occurs, a value of 0 is set as lic_flag and encoding is performed without applying LIC processing. On the other hand, a decoding device decodes lic_flag described in a stream, and switches whether or not to apply LIC processing depending on the value and performs decoding.

[0121] Another method of determining whether to apply LIC processing is, for example, a method of determining according to whether LIC processing is applied to surrounding blocks.As a specific example, when the block to be coded is in merge mode, determine whether the surrounding coded blocks selected when deriving MV in merge mode processing have been coded by applying LIC processing, and switch whether to apply LIC processing according to the result and perform coding.In addition, in this example, the process in decoding is exactly the same.

[0122] [Overview of the Decryption Device] Next, a description will be given of an overview of a decoding device capable of decoding the coded signal (coded bit stream) output from the above coding device 100. Fig. 10 is a block diagram showing a functional configuration of a decoding device 200 according to the first embodiment. The decoding device 200 is a video / image decoding device that decodes a video / image on a block-by-block basis.

[0123] As shown in FIG. 10, the decoding device 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transform unit 206, an addition unit 208, a block memory 210, a loop filter unit 212, a frame memory 214, an intra prediction unit 216, an inter prediction unit 218, and a prediction control unit 220.

[0124] The decoding device 200 is realized by, for example, a general-purpose processor and a memory. In this case, when the software program stored in the memory is executed by the processor, the processor functions as the entropy decoding unit 202, the inverse quantization unit 204, the inverse transform unit 206, the addition unit 208, the loop filter unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220. The decoding device 200 may also be realized as one or more dedicated electronic circuits corresponding to the entropy decoding unit 202, the inverse quantization unit 204, the inverse transform unit 206, the addition unit 208, the loop filter unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220.

[0125] Each component included in the decoding device 200 will be described below.

[0126] [Entropy Decoding Part] The entropy decoding unit 202 entropy decodes the coded bit stream. Specifically, the entropy decoding unit 202 arithmetically decodes the coded bit stream into a binary signal, for example. The entropy decoding unit 202 then debinarizes the binary signal. As a result, the entropy decoding unit 202 outputs quantized coefficients to the inverse quantization unit 204 on a block-by-block basis.

[0127] [Dequantization section] The inverse quantization unit 204 inverse quantizes the quantized coefficients of a block to be decoded (hereinafter, referred to as a current block) that is input from the entropy decoding unit 202. Specifically, the inverse quantization unit 204 inverse quantizes each quantized coefficient of the current block based on a quantization parameter corresponding to the quantized coefficient. Then, the inverse quantization unit 204 outputs the inverse quantized quantized coefficients (i.e., transform coefficients) of the current block to the inverse transform unit 206.

[0128] [Inverse conversion section] The inverse transform unit 206 restores the prediction error by inverse transforming the transform coefficients input from the inverse quantization unit 204 .

[0129] For example, if the information interpreted from the encoded bitstream indicates that EMT or AMT is to be applied (e.g., the AMT flag is true), the inverse transform unit 206 inverse transforms the transform coefficients of the current block based on the interpreted information indicating the transform type.

[0130] Also for example, if the information interpreted from the coded bitstream indicates to apply NSST, then inverse transform unit 206 applies an inverse re-transform to the transform coefficients.

[0131] [Addition section] The adder 208 reconstructs the current block by adding the prediction error, which is an input from the inverse transformer 206, and the prediction sample, which is an input from the prediction control unit 220. The adder 208 then outputs the reconstructed block to the block memory 210 and the loop filter unit 212.

[0132] [Block memory] The block memory 210 is a storage unit for storing blocks that are referenced in intra prediction and are in a picture to be decoded (hereinafter, referred to as a current picture). Specifically, the block memory 210 stores the reconstructed block output from the adder 208.

[0133] [Loop filter section] The loop filter unit 212 applies a loop filter to the block reconstructed by the adder unit 208, and outputs the filtered reconstructed block to a frame memory 214, a display device, or the like.

[0134] If the information indicating ALF on / off read from the encoded bitstream indicates ALF on, one filter is selected from among multiple filters based on the local gradient direction and activity, and the selected filter is applied to the reconstructed block.

[0135] [Frame memory] The frame memory 214 is a storage unit for storing reference pictures used in inter prediction, and may be called a frame buffer. Specifically, the frame memory 214 stores the reconstructed blocks filtered by the loop filter unit 212.

[0136] [Intra prediction section] The intra prediction unit 216 generates a prediction signal (intra prediction signal) by performing intra prediction with reference to a block in the current picture stored in the block memory 210 based on an intra prediction mode interpreted from the encoded bit stream. Specifically, the intra prediction unit 216 generates an intra prediction signal by performing intra prediction with reference to samples (e.g., luminance values, chrominance values) of blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control unit 220.

[0137] Note that, when an intra prediction mode that references a luminance block in intra prediction of a chrominance block is selected, the intra prediction unit 216 may predict the chrominance component of the current block based on the luminance component of the current block.

[0138] Furthermore, when information interpreted from the encoded bitstream indicates the application of PDPC, the intra prediction unit 216 corrects pixel values ​​after intra prediction based on the gradients of reference pixels in the horizontal / vertical directions.

[0139] [Inter prediction section] The inter prediction unit 218 predicts the current block by referring to a reference picture stored in the frame memory 214. The prediction is performed in units of the current block or sub-blocks (e.g., 4x4 blocks) in the current block. For example, the inter prediction unit 218 generates an inter prediction signal of the current block or sub-block by performing motion compensation using motion information (e.g., motion vectors) interpreted from the encoded bitstream, and outputs the inter prediction signal to the prediction control unit 220.

[0140] In addition, when the information interpreted from the encoded bitstream indicates that the OBMC mode is to be applied, the inter prediction unit 218 generates an inter prediction signal using not only the motion information of the current block obtained by motion search, but also the motion information of adjacent blocks.

[0141] Also, if the information interpreted from the encoded bitstream indicates that the FRUC mode is applied, the inter prediction unit 218 derives motion information by performing motion search according to the pattern matching method (bilateral matching or template matching) interpreted from the encoded bitstream. Then, the inter prediction unit 218 performs motion compensation using the derived motion information.

[0142] In addition, when the BIO mode is applied, the inter prediction unit 218 derives a motion vector based on a model assuming uniform linear motion. In addition, when information interpreted from the encoded bitstream indicates that an affine motion compensation prediction mode is applied, the inter prediction unit 218 derives a motion vector on a sub-block basis based on the motion vectors of multiple adjacent blocks.

[0143] [Predictive control unit] The prediction control unit 220 selects either the intra-prediction signal or the inter-prediction signal, and outputs the selected signal to the addition unit 208 as a prediction signal.

[0144] (Embodiment 2) The encoding process and the decoding process according to the second embodiment will be specifically described with reference to FIGS. 11 and 12, and the encoding device and the decoding device according to the second embodiment will be specifically described with reference to FIGS. 15 and 16.

[0145] [Encoding process] FIG. 11 shows a video encoding process according to the second embodiment.

[0146] First, in step S1001, a first parameter is written into a bitstream, the first parameter identifying a partition mode from among a plurality of partition modes for dividing a first block into a plurality of sub-blocks. When a partition mode is used, the block is divided into a plurality of sub-blocks. When different partition modes are used, the block is divided into a plurality of sub-blocks with different shapes, different heights, or different widths.

[0147] FIG. 28 shows an example of a partition mode for dividing a block of NxN pixels in the second embodiment. In FIG. 28, (a) to (h) show partition modes different from each other. As shown in FIG. 28, when partition mode (a) is used, a block of NxN pixels (e.g., 16x16 pixels, where the value of "N" can be any value that is an integer multiple of 4 between 8 and 128) is divided into two sub-blocks of N / 2xN pixels (e.g., 8x16 pixels). When partition mode (b) is used, the block of NxN pixels is divided into a sub-block of N / 4xN pixels (e.g., 4x16 pixels) and a sub-block of 3N / 4xN pixels (e.g., 12x16 pixels). When partition mode (c) is used, the block of NxN pixels is divided into a sub-block of 3N / 4xN pixels (e.g., 12x16 pixels) and a sub-block of N / 4xN pixels (e.g., 4x16 pixels). Using partition mode (d), a block of NxN pixels is divided into a subblock of (N / 4)xN pixels (e.g., 4x16 pixels), a subblock of N / 2xN pixels (e.g., 8x16 pixels), and a subblock of N / 4xN pixels (e.g., 4x16 pixels). Using partition mode (e), a block of NxN pixels is divided into two subblocks of NxN / 2 pixels (e.g., 16x8 pixels). Using partition mode (f), a block of NxN pixels is divided into a subblock of NxN / 4 pixels (e.g., 16x4 pixels) and a subblock of Nx3N / 4 pixels (e.g., 16x12 pixels). Using partition mode (g), a block of NxN pixels is divided into a subblock of Nx3N / 4 pixels (e.g., 16x12 pixels) and a subblock of NxN / 4 pixels (e.g., 16x4 pixels). Using partition mode (h), a block of NxN pixels is divided into a sub-block of NxN / 4 pixels (e.g., 16x4 pixels), a sub-block of NxN / 2 pixels (e.g., 16x8 pixels), and a sub-block of NxN / 4 pixels (e.g., 16x4 pixels).

[0148] Next, in step S1002, it is determined whether the first parameter identifies a first partition mode.

[0149] Next, in step S1003, it is determined whether to select the second partition mode as a candidate for dividing the second block based at least on a determination of whether the first parameter identifies the first partition mode.

[0150] Two different partition mode sets may divide a block into subblocks of the same shape and size. For example, as shown in FIG. 31A, subblocks (1b) and (2c) have the same shape and size. One partition mode set may include at least two partition modes. For example, as shown in FIG. 31A, (1a) and (1b), one partition mode set may include ternary tree vertical division followed by binary tree vertical division of the middle subblock and no division of the other subblocks. And for example, as shown in FIG. 31A, (2a), (2b) and (2c), another partition mode set may include binary tree vertical division followed by binary tree vertical division of both subblocks. Both partition mode sets result in subblocks of the same shape and size.

[0151] When choosing between two partition mode sets that divide a block into sub-blocks of the same shape and size, and that have different numbers of bins or bits when encoded into a bitstream, the partition mode set with the fewer number of bins or bits is selected, where the number of bins and the number of bits correspond to the amount of code.

[0152] When choosing between two partition mode sets that divide a block into sub-blocks of the same shape and size and that have the same number of bins or bits when encoded into a bitstream, the partition mode set that appears first in a predetermined order of the multiple partition mode sets is selected, which may be, for example, based on the number of partition modes in each partition mode set.

[0153] 31A and 31B are diagrams showing an example of dividing a block into sub-blocks using a partition mode set with a smaller number of bins in partition mode encoding. In this example, when the left NxN pixel block is vertically divided into two sub-blocks, the second partition mode for the right NxN pixel block is not selected in step (2c). This is because, in the partition mode encoding method of FIG. 31B, the second partition mode set (2a, 2b, 2c) requires more bins for partition mode encoding compared to the first partition mode set (1a, 1b).

[0154] 32A to 32C are diagrams illustrating an example of dividing a block into sub-blocks using a partition mode set that appears first in a predetermined order of multiple partition mode sets. In this example, when a block of 2NxN / 2 pixels is vertically divided into three sub-blocks, the second partition mode for the lower block of 2NxN / 2 pixels is not selected in step (2c). This is because, in the partition mode encoding method of FIG. 32B, the second partition mode set (2a, 2b, 2c) has the same number of bins as the first partition mode set (1a, 1b, 1c, 1d) and appears after the first partition mode set (1a, 1b, 1c, 1d) in the predetermined order of partition mode sets illustrated in FIG. 32C. The predetermined order of multiple partition mode sets can be fixed or signaled in the bitstream.

[0155] Figure 20 shows an example in which the second partition mode is not selected for dividing a block of 2NxN pixels, as shown in step (2c) in the second embodiment. As shown in Figure 20, the first division method (i) can be used to equally divide a block of 2Nx2N pixels (e.g., 16x16 pixels) into four sub-blocks of NxN pixels (e.g., 8x8 pixels), as shown in step (1a). Also, the second division method (ii) can be used to equally divide a block of 2Nx2N pixels horizontally into two sub-blocks of 2NxN pixels (e.g., 16x8 pixels), as shown in step (2a). Here, in the second division method (ii), if the upper 2NxN pixel block (first block) is vertically divided into two NxN pixel sub-blocks by the first partition mode as in step (2b), the second partition mode of vertically dividing the lower 2NxN pixel block (second block) into two NxN pixel sub-blocks in step (2c) is not selected as a possible partition mode candidate, because it generates the same sub-block size as the sub-block size obtained by the quadrant division in the first division method (i).

[0156] As described above, in FIG. 20, if the first partition mode is used to equally divide the first block into two sub-blocks vertically, and the second partition mode is used to equally divide the second block vertically adjacent to the first block into two sub-blocks vertically, then the second partition mode is not selected as a candidate.

[0157] FIG. 21 shows an example in which the second partition mode is not selected for dividing a block of Nx2N pixels as shown in step (2c) in the second embodiment. As shown in FIG. 21, the first division method (i) can be used to equally divide a block of 2Nx2N pixels into four sub-blocks of NxN pixels as shown in step (1a). Also, the second division method (ii) can be used to equally divide a block of 2Nx2N pixels vertically into two sub-blocks of 2NxN pixels (e.g., 8x16 pixels) as shown in step (2a). In the second division method (ii), when the block of Nx2N pixels (first block) on the left side is divided horizontally into two sub-blocks of NxN pixels by the first partition mode as shown in step (2b), the second partition mode that divides the block of Nx2N pixels (second block) on the right side horizontally into two sub-blocks of NxN pixels in step (2c) is not selected as a possible candidate partition mode. This is because the sub-block size generated is the same as the sub-block size obtained by the four-division in the first division method (i).

[0158] As described above, in FIG. 21, if the first partition mode is used to divide the first block equally into two sub-blocks horizontally, and the second partition mode is used to divide the second block horizontally adjacent to the first block equally into two sub-blocks horizontally, then the second partition mode is not selected as a candidate.

[0159] FIG. 40 shows an example of dividing the 4Nx2N partition in FIG. 20 into three partitions in a ratio of 1:2:1, such as Nx2N, 2Nx2N, and Nx2N. Here, when the upper block is divided into three, a partition mode that divides the lower block into three partitions in a ratio of 1:2:1 is not selected as a possible partition mode candidate. The division into three partitions may be in a ratio different from 1:2:1. Furthermore, the partition may be divided into three or more partitions, or may be divided into two partitions, or may have a ratio different from 1:1, such as 1:2 or 1:3. FIG. 40 shows an example of dividing the partition horizontally first, but similar constraints can be applied when dividing the partition vertically first.

[0160] 41 and 42 show an example of applying a similar constraint when the first block is rectangular.

[0161] FIG. 43 shows a second example of constraint when dividing a square vertically into thirds and then horizontally into two equal parts. When the constraint of FIG. 43 is applied, a partition mode in which the lower block of 4Nx2N is divided into three parts at a ratio of 1:2:1 in FIG. 40 can be selected. Information indicating whether the constraint of FIG. 40 or the constraint of FIG. 43 is to be applied may be coded separately in header information or the like. Alternatively, a constraint may be applied so that the amount of code of the information indicating the partition is reduced. For example, if the amount of code of the information indicating the partition in case 1 and case 2 is equal to or less than the amount of code, the division in case 1 is valid and the division in case 2 is invalid. That is, the constraint of FIG. 43 is applied.

[0162] (Case 1) (1) After dividing the square horizontally into two, (2) each of the top and bottom rectangular blocks is divided vertically into three: (1) Direction information: 1 bit, number of divisions information: 1 bit, (2) (Direction information: 1 bit, number of divisions information: 1 bit) x 2, totaling 6 bits. (Case 2) (1) After dividing the square vertically, (2) each of the left, center, and right rectangular blocks is divided horizontally into two: (1) Direction information: 1 bit, number of divisions information: 1 bit, (2) (Direction information: 1 bit, number of divisions information: 1 bit) x 3, totaling 8 bits.

[0163] Alternatively, the optimal partition may be determined while selecting a partition mode in a predetermined order during encoding. For example, it is possible to first try 2-partitioning, then 3-partitioning or 4-partitioning (dividing into 2 equal parts horizontally and vertically), and so on. In this case, before the trial of 3-partitioning as in FIG. 43, a trial starting from 2-partitioning as in the example of FIG. 40 has already been performed. Therefore, in the trial starting from 2-partitioning, a partition that divides equally horizontally and further divides the upper and lower blocks into 3 equal parts vertically has already been tried, so the constraint in FIG. 43 is applied. In this way, the constraint method to be selected may be determined based on a predetermined encoding method.

[0164] Fig. 44 shows an example in which the second partition mode restricts the selectable partition modes for the same direction as the first partition mode. Here, the first partition mode is vertical 3-partition, and in this case, 2-partition cannot be selected as the second partition mode. On the other hand, 2-partition can be selected for the vertical direction, which is different from the first partition mode (Fig. 45).

[0165] FIG. 22 shows an example in which the second partition mode is not selected for dividing a block of NxN pixels as shown in step (2c) in the second embodiment. As shown in FIG. 22, using the first division method (i), a block of 2NxN pixels (e.g., 16x8 pixels, where the value of "N" can be any value that is an integer multiple of 4 from 8 to 128) can be divided vertically into a sub-block of N / 2xN pixels, a sub-block of NxN pixels, and a sub-block of N / 2xN pixels (e.g., a sub-block of 4x8 pixels, a sub-block of 8x8 pixels, and a sub-block of 4x8 pixels) as shown in step (1a). Also, using the second division method (ii), a block of 2NxN pixels can be divided into two sub-blocks of NxN pixels as shown in step (2a). In the first division method (i), the central block of NxN pixels can be vertically divided into two subblocks of N / 2xN pixels (e.g., 4x8 pixels) in step (1b). In the second division method (ii), if the left block of NxN pixels (first block) is vertically divided into two subblocks of N / 2xN pixels as in step (2b), the partition mode of vertically dividing the block of NxN pixels (second block) on the right side into two subblocks of N / 2xN pixels in step (2c) is not selected as a possible partition mode candidate. This is because the subblock size is the same as that obtained by the first division method (i), that is, four subblocks of N / 2xN pixels are generated.

[0166] As described above, in FIG. 22, if the first partition mode is used to equally divide the first block into two sub-blocks vertically, and the second partition mode is used to equally divide the second block horizontally adjacent to the first block into two sub-blocks vertically, then the second partition mode is not selected as a candidate.

[0167] FIG. 23 shows an example in which the second partition mode is not selected for dividing a block of NxN pixels as shown in step (2c) in the second embodiment. As shown in FIG. 23, using the first division method (i), Nx2N pixels (e.g., 8x16 pixels, where the value of "N" can be any value that is an integer multiple of 4 between 8 and 128) can be divided into a subblock of NxN / 2 pixels, a subblock of NxN pixels, and a subblock of NxN / 2 pixels (e.g., a subblock of 8x4 pixels, a subblock of 8x8 pixels, and a subblock of 8x4 pixels) as shown in step (1a). Also, using the second division method, it can be divided into two subblocks of NxN pixels as shown in step (2a). In the first division method (i), the central block of NxN pixels can be divided into two subblocks of NxN / 2 pixels as shown in step (1b). In the second division method (ii), when the upper NxN pixel block (first block) is horizontally divided into two NxN / 2 pixel sub-blocks as in step (2b), the partition mode of horizontally dividing the lower NxN pixel block (second block) into two NxN / 2 pixel sub-blocks in step (2c) is not selected as a possible partition mode candidate, because it generates the same sub-block size as the sub-block size obtained by the first division method (i), i.e., four NxN / 2 pixel sub-blocks.

[0168] As described above, in FIG. 23, if the first partition mode is used to divide the first block equally into two sub-blocks horizontally, and the second partition mode is used to divide the second block vertically adjacent to the first block equally into two sub-blocks horizontally, then the second partition mode is not selected as a candidate.

[0169] If it is determined that the second partition mode is selected as a candidate for dividing the second block (N in S1003), then in step S1004, a partition mode is selected from a plurality of partition modes including the second partition mode as a candidate. In step S1005, a second parameter indicating the selection result is written to the bitstream.

[0170] If it is determined that the second partition mode is not selected as a candidate for dividing the second block (Y in S1003), then in step S1006, a partition mode different from the second partition mode is selected for dividing the second block, where the selected partition mode divides the block into sub-blocks having a different shape or a different size compared to the sub-blocks generated by the second partition mode.

[0171] FIG. 24 shows an example of dividing a block of 2NxN pixels using the selected partition mode when the second partition mode is not selected as shown in step (3) in the second embodiment. As shown in FIG. 24, the selected partition mode can divide the current block of 2NxN pixels (the bottom block in this example) into three sub-blocks as shown in (c) and (f) of FIG. 24. The three sub-blocks may have different sizes. For example, in the three sub-blocks, the large sub-block may have twice the width / height of the small sub-block. Also, for example, the selected partition mode can divide the current block into two sub-blocks of different sizes (asymmetric binary tree) as shown in (a), (b), (d) and (e) of FIG. 24. For example, when an asymmetric binary tree is used, the large sub-block may have three times the width / height of the small sub-block.

[0172] FIG. 25 shows an example of dividing a block of Nx2N pixels using the selected partition mode when the second partition mode is not selected as shown in step (3) in the second embodiment. As shown in FIG. 25, the selected partition mode can divide the current block of Nx2N pixels (the right block in this example) into three sub-blocks as shown in (c) and (f) of FIG. 25. The three sub-blocks may have different sizes. For example, in the three sub-blocks, the large sub-block may have twice the width / height of the small sub-block. Also, for example, the selected partition mode can divide the current block into two sub-blocks of different sizes (asymmetric binary tree) as shown in (a), (b), (d) and (e) of FIG. 25. For example, when an asymmetric binary tree is used, the large sub-block may have three times the width / height of the small sub-block.

[0173] FIG. 26 shows an example of dividing a block of NxN pixels using a partition mode selected when the second partition mode is not selected as shown in step (3) in the second embodiment. As shown in FIG. 26, in step (1), a block of 2NxN pixels is divided vertically into two sub-blocks of NxN pixels, and in step (2), a block of NxN pixels on the left side is divided vertically into two sub-blocks of N / 2xN pixels. In step (3), the selected partition mode for the current block of NxN pixels (the left block in this example) can be used to divide the current block into three sub-blocks as shown in (c) and (f) of FIG. 26. The sizes of the three sub-blocks may be different. For example, in the three sub-blocks, the large sub-block may have twice the width / height of the small sub-block. Also, for example, the selected partition mode can divide the current block into two sub-blocks of different sizes (asymmetric binary tree) as shown in (a), (b), (d), and (e) of FIG. 26. For example, if an asymmetric binary tree is used, the large sub-blocks may have three times the width / height of the small sub-blocks.

[0174] FIG. 27 shows an example of dividing a block of NxN pixels using a partition mode selected when the second partition mode is not selected as shown in step (3) in the second embodiment. As shown in FIG. 27, in step (1), a block of Nx2N pixels is divided horizontally into two sub-blocks of NxN pixels, and in step (2), the upper block of NxN pixels is divided horizontally into two sub-blocks of NxN / 2 pixels. In step (3), the selected partition mode for the current block of NxN pixels (the lower block in this example) can be used to divide the current block into three sub-blocks as shown in (c) and (f) of FIG. 27. The sizes of the three sub-blocks may be different. For example, in the three sub-blocks, the large sub-block may have twice the width / height of the small sub-block. Also, for example, the selected partition mode can divide the current block into two sub-blocks of different sizes (asymmetric binary tree) as shown in (a), (b), (d), and (e) of FIG. 27. For example, if an asymmetric binary tree is used, the large sub-blocks may have three times the width / height of the small sub-blocks.

[0175] Figure 17 illustrates possible locations of the first parameter in a compressed video stream. As illustrated in Figure 17, the first parameter may be located in a video parameter set, a sequence parameter set, a picture parameter set, a slice header, or a coding tree unit. The first parameter may indicate how to divide the block into multiple sub-blocks. For example, the first parameter may include a flag indicating whether to divide the block horizontally or vertically. The first parameter may also include a parameter indicating whether to divide the block into two or more sub-blocks.

[0176] FIG. 18 illustrates possible locations of the second parameter in the compressed video stream. As shown in FIG. 18, the second parameter may be located in a video parameter set, a sequence parameter set, a picture parameter set, a slice header, or a coding tree unit. The second parameter may indicate how to divide the block into multiple sub-blocks. For example, the second parameter may include a flag indicating whether to divide the block horizontally or vertically. The second parameter may also include a parameter indicating whether to divide the block into two or more sub-blocks. The second parameter is located following the first parameter in the bitstream, as shown in FIG. 19.

[0177] The first block and the second block are different blocks. The first block and the second block may be included in the same frame. For example, the first block may be an adjacent block above the second block. Also, for example, the first block may be an adjacent block to the left of the second block.

[0178] In step S1007, the second block is divided into sub-blocks using the selected partition mode. In step S1008, the divided blocks are coded.

[0179] [Encoding device] FIG. 15 is a block diagram showing a structure of a video / image coding device according to the second or third embodiment. In FIG.

[0180] The video encoding device 5000 is a device for encoding an input video / image for each block to generate an encoded output bitstream. As shown in Fig. 15, the video encoding device 5000 includes a transform unit 5001, a quantization unit 5002, an inverse quantization unit 5003, an inverse transform unit 5004, a block memory 5005, a frame memory 5006, an intra prediction unit 5007, an inter prediction unit 5008, an entropy encoding unit 5009, and a block division determination unit 5010.

[0181] An input image is input to an adder, and an added value is output to a conversion unit 5001. The conversion unit 5001 converts the added value into a frequency coefficient based on a block partition mode derived by a block partition determination unit 5010, and outputs the frequency coefficient to a quantization unit 5002. The block partition mode may be associated with a block partition mode, a block partition type, or a block partition direction. The quantization unit 5002 quantizes the input quantization coefficient, and outputs the quantization value to an inverse quantization unit 5003 and an entropy coding unit 5009.

[0182] The inverse quantization unit 5003 inversely quantizes the quantized value output from the quantization unit 5002, and outputs the frequency coefficients to the inverse transform unit 5004. The inverse transform unit 5004 performs an inverse frequency transform on the frequency coefficients based on the block partition mode derived by the block partition determination unit 5010, converts the frequency coefficients into sample values ​​of a bit stream, and outputs the sample values ​​to an adder.

[0183] The adder adds the sample values ​​of the bitstream output from the inverse transform unit 5004 to the predicted video / image values ​​output from the intra / inter prediction units 5007 and 5008, and outputs the sum to the block memory 5005 or the frame memory 5006 for further prediction. The block partition determination unit 5010 collects block information from the block memory 5005 or the frame memory 5006, and derives a block partition mode and parameters related to the block partition mode. Using the derived block partition mode, the block is divided into multiple sub-blocks. The intra / inter prediction units 5007 and 5008 search among the video / images stored in the block memory 5005 or the video / images in the frame memory 5006 reconstructed in the block partition mode derived by the block partition determination unit 5010, and estimate, for example, a video / image area that is most similar to the input video / image to be predicted.

[0184] The entropy coding unit 5009 codes the quantized value output from the quantization unit 5002, codes the parameters from the block division determination unit 5010, and outputs a bit stream.

[0185] [Decryption process] FIG. 12 shows a video decoding process according to the second embodiment.

[0186] First, in step S2001, a first parameter is read from the bitstream, the first parameter identifying a partition mode for dividing a first block into sub-blocks among a plurality of partition modes, whereby the block is divided into sub-blocks, and different partition modes are used to divide the block into sub-blocks having different shapes, different heights or different widths.

[0187] FIG. 28 shows an example of a partition mode for dividing a block of NxN pixels in the second embodiment. In FIG. 28, (a) to (h) show partition modes different from each other. As shown in FIG. 28, when partition mode (a) is used, a block of NxN pixels (e.g., 16x16 pixels, where the value of "N" can be any value that is an integer multiple of 4 between 8 and 128) is divided into two sub-blocks of N / 2xN pixels (e.g., 8x16 pixels). When partition mode (b) is used, the block of NxN pixels is divided into a sub-block of N / 4xN pixels (e.g., 4x16 pixels) and a sub-block of 3N / 4xN pixels (e.g., 12x16 pixels). When partition mode (c) is used, the block of NxN pixels is divided into a sub-block of 3N / 4xN pixels (e.g., 12x16 pixels) and a sub-block of N / 4xN pixels (e.g., 4x16 pixels). Using partition mode (d), a block of NxN pixels is divided into a subblock of (N / 4)xN pixels (e.g., 4x16 pixels), a subblock of N / 2xN pixels (e.g., 8x16 pixels), and a subblock of N / 4xN pixels (e.g., 4x16 pixels). Using partition mode (e), a block of NxN pixels is divided into two subblocks of NxN / 2 pixels (e.g., 16x8 pixels). Using partition mode (f), a block of NxN pixels is divided into a subblock of NxN / 4 pixels (e.g., 16x4 pixels) and a subblock of Nx3N / 4 pixels (e.g., 16x12 pixels). Using partition mode (g), a block of NxN pixels is divided into a subblock of Nx3N / 4 pixels (e.g., 16x12 pixels) and a subblock of NxN / 4 pixels (e.g., 16x4 pixels). Using partition mode (h), a block of NxN pixels is divided into a sub-block of NxN / 4 pixels (e.g., 16x4 pixels), a sub-block of NxN / 2 pixels (e.g., 16x8 pixels), and a sub-block of NxN / 4 pixels (e.g., 16x4 pixels).

[0188] Next, in step S2002, a determination is made whether the first parameter identifies a first partition mode.

[0189] Next, in step S2003, it is determined whether to select the second partition mode as a candidate for dividing the second block based at least on a determination of whether the first parameter identifies the first partition mode.

[0190] Two different partition mode sets may divide a block into subblocks of the same shape and size. For example, as shown in FIG. 31A, subblocks (1b) and (2c) have the same shape and size. One partition mode set may include at least two partition modes. For example, as shown in FIG. 31A, (1a) and (1b), one partition mode set may include ternary tree vertical division followed by binary tree vertical division of the middle subblock and no division of the other subblocks. And for example, as shown in FIG. 31A, (2a), (2b) and (2c), another partition mode set may include binary tree vertical division followed by binary tree vertical division of both subblocks. Both partition mode sets result in subblocks of the same shape and size.

[0191] When choosing between two partition mode sets that divide a block into sub-blocks of the same shape and size, but that have different numbers of bins or bits when encoded into the bitstream, the partition mode set with the fewer number of bins or bits is selected.

[0192] When choosing between two partition mode sets that divide a block into sub-blocks of the same shape and size and that have the same number of bins or bits when encoded into a bitstream, the partition mode set that appears first in a predetermined order of the multiple partition mode sets is selected, which may be, for example, based on the number of partition modes in each partition mode set.

[0193] 31A and 31B are diagrams showing an example of dividing a block into sub-blocks using a partition mode set with a smaller number of bins in partition mode encoding. In this example, when the left NxN pixel block is vertically divided into two sub-blocks, the second partition mode for the right NxN pixel block is not selected in step (2c). This is because, in the partition mode encoding method of FIG. 31B, the second partition mode set (2a, 2b, 2c) requires more bins for partition mode encoding compared to the first partition mode set (1a, 1b).

[0194] FIG. 32A is a diagram showing an example of dividing a block into sub-blocks using a partition mode set that appears first in a predetermined order of multiple partition mode sets. In this example, when a block of 2NxN / 2 pixels is vertically divided into three sub-blocks, the second partition mode for the lower block of 2NxN / 2 pixels is not selected in step (2c). This is because, in the partition mode encoding method of FIG. 32B, the second partition mode set (2a, 2b, 2c) has the same number of bins as the first partition mode set (1a, 1b, 1c, 1d) and appears after the first partition mode set (1a, 1b, 1c, 1d) in the predetermined order of partition mode sets shown in FIG. 32C. The predetermined order of multiple partition mode sets can be fixed or signaled in the bitstream.

[0195] Figure 20 shows an example in which the second partition mode is not selected for dividing a block of 2NxN pixels, as shown in step (2c) in the second embodiment. As shown in Figure 20, the first division method (i) can be used to equally divide a block of 2Nx2N pixels (e.g., 16x16 pixels) into four sub-blocks of NxN pixels (e.g., 8x8 pixels), as shown in step (1a). Also, the second division method (ii) can be used to equally divide a block of 2Nx2N pixels horizontally into two sub-blocks of 2NxN pixels (e.g., 16x8 pixels), as shown in step (2a). In the second division method (ii), if the first partition mode divides the upper 2NxN pixel block (first block) vertically into two NxN pixel sub-blocks in step (2b), the second partition mode dividing the lower 2NxN pixel block (second block) vertically into two NxN pixel sub-blocks in step (2c) is not selected as a possible partition mode candidate because it would generate the same sub-block size as the sub-block size obtained by the quadrant division in the first division method (i).

[0196] As described above, in FIG. 20, if the first partition mode is used to equally divide the first block into two sub-blocks vertically, and the second partition mode is used to equally divide the second block vertically adjacent to the first block into two sub-blocks vertically, then the second partition mode is not selected as a candidate.

[0197] FIG. 21 shows an example in which the second partition mode is not selected for dividing a block of Nx2N pixels as shown in step (2c) in the second embodiment. As shown in FIG. 21, the first division method (i) can be used to equally divide a block of 2Nx2N pixels into four sub-blocks of NxN pixels as shown in step (1a). Also, the second division method (ii) can be used to equally divide a block of 2Nx2N pixels vertically into two sub-blocks of 2NxN pixels (e.g., 8x16 pixels) as shown in step (2a). In the second division method (ii), when the block of Nx2N pixels (first block) on the left side is divided horizontally into two sub-blocks of NxN pixels by the first partition mode as shown in step (2b), the second partition mode that divides the block of Nx2N pixels (second block) on the right side horizontally into two sub-blocks of NxN pixels in step (2c) is not selected as a possible candidate partition mode. This is because the sub-block size generated is the same as the sub-block size obtained by the four-division in the first division method (i).

[0198] As described above, in FIG. 21, if the first partition mode is used to divide the first block equally into two sub-blocks horizontally, and the second partition mode is used to divide the second block horizontally adjacent to the first block equally into two sub-blocks horizontally, then the second partition mode is not selected as a candidate.

[0199] FIG. 22 shows an example in which the second partition mode is not selected for dividing a block of NxN pixels as shown in step (2c) in the second embodiment. As shown in FIG. 22, using the first division method (i), a block of 2NxN pixels (e.g., 16x8 pixels, where the value of "N" can be any value that is an integer multiple of 4 from 8 to 128) can be divided vertically into a sub-block of N / 2xN pixels, a sub-block of NxN pixels, and a sub-block of N / 2xN pixels (e.g., a sub-block of 4x8 pixels, a sub-block of 8x8 pixels, and a sub-block of 4x8 pixels) as shown in step (1a). Also, using the second division method (ii), a block of 2NxN pixels can be divided into two sub-blocks of NxN pixels as shown in step (2a). In the first division method (i), the central block of NxN pixels can be vertically divided into two subblocks of N / 2xN pixels (e.g., 4x8 pixels) in step (1b). In the second division method (ii), if the left block of NxN pixels (first block) is vertically divided into two subblocks of N / 2xN pixels as in step (2b), the partition mode of vertically dividing the block of NxN pixels (second block) on the right side into two subblocks of N / 2xN pixels in step (2c) is not selected as a possible partition mode candidate. This is because the subblock size is the same as that obtained by the first division method (i), that is, four subblocks of N / 2xN pixels are generated.

[0200] As described above, in FIG. 22, if the first partition mode is used to equally divide the first block into two sub-blocks vertically, and the second partition mode is used to equally divide the second block horizontally adjacent to the first block into two sub-blocks vertically, then the second partition mode is not selected as a candidate.

[0201] FIG. 23 shows an example in which the second partition mode is not selected for dividing a block of NxN pixels as shown in step (2c) in the second embodiment. As shown in FIG. 23, using the first division method (i), Nx2N pixels (e.g., 8x16 pixels, where the value of "N" can be any value that is an integer multiple of 4 between 8 and 128) can be divided into a subblock of NxN / 2 pixels, a subblock of NxN pixels, and a subblock of NxN / 2 pixels (e.g., a subblock of 8x4 pixels, a subblock of 8x8 pixels, and a subblock of 8x4 pixels) as shown in step (1a). Also, using the second division method, it can be divided into two subblocks of NxN pixels as shown in step (2a). In the first division method (i), the central block of NxN pixels can be divided into two subblocks of NxN / 2 pixels as shown in step (1b). In the second division method (ii), when the upper NxN pixel block (first block) is horizontally divided into two NxN / 2 pixel sub-blocks as in step (2b), the partition mode of horizontally dividing the lower NxN pixel block (second block) into two NxN / 2 pixel sub-blocks in step (2c) is not selected as a possible partition mode candidate, because it generates the same sub-block size as the sub-block size obtained by the first division method (i), i.e., four NxN / 2 pixel sub-blocks.

[0202] As described above, in FIG. 23, if the first partition mode is used to divide the first block equally into two sub-blocks horizontally, and the second partition mode is used to divide the second block vertically adjacent to the first block equally into two sub-blocks horizontally, then the second partition mode is not selected as a candidate.

[0203] If it is determined that the second partition mode is to be selected as a candidate for dividing the second block (N in S2003), in step S2004, the second parameter is decoded from the bitstream, and a partition mode is selected from a plurality of partition modes including the second partition mode as a candidate.

[0204] If it is determined that the second partition mode is not selected as a candidate for dividing the second block (Y in S2003), then in step S2005, a partition mode different from the second partition mode is selected for dividing the second block, where the selected partition mode divides the block into sub-blocks having a different shape or a different size compared to the sub-blocks generated by the second partition mode.

[0205] FIG. 24 shows an example of dividing a block of 2NxN pixels using the selected partition mode when the second partition mode is not selected as shown in step (3) in the second embodiment. As shown in FIG. 24, the selected partition mode can divide the current block of 2NxN pixels (the bottom block in this example) into three sub-blocks as shown in (c) and (f) of FIG. 24. The three sub-blocks may have different sizes. For example, in the three sub-blocks, the large sub-block may have twice the width / height of the small sub-block. Also, for example, the selected partition mode can divide the current block into two sub-blocks of different sizes (asymmetric binary tree) as shown in (a), (b), (d) and (e) of FIG. 24. For example, when an asymmetric binary tree is used, the large sub-block may have three times the width / height of the small sub-block.

[0206] FIG. 25 shows an example of dividing a block of Nx2N pixels using the selected partition mode when the second partition mode is not selected as shown in step (3) in the second embodiment. As shown in FIG. 25, the selected partition mode can divide the current block of Nx2N pixels (the right block in this example) into three sub-blocks as shown in (c) and (f) of FIG. 25. The three sub-blocks may have different sizes. For example, in the three sub-blocks, the large sub-block may have twice the width / height of the small sub-block. Also, for example, the selected partition mode can divide the current block into two sub-blocks of different sizes (asymmetric binary tree) as shown in (a), (b), (d) and (e) of FIG. 25. For example, when an asymmetric binary tree is used, the large sub-block may have three times the width / height of the small sub-block.

[0207] FIG. 26 shows an example of dividing a block of NxN pixels using a partition mode selected when the second partition mode is not selected as shown in step (3) in the second embodiment. As shown in FIG. 26, in step (1), a block of 2NxN pixels is divided vertically into two sub-blocks of NxN pixels, and in step (2), a block of NxN pixels on the left side is divided vertically into two sub-blocks of N / 2xN pixels. In step (3), the selected partition mode for the current block of NxN pixels (the left block in this example) can be used to divide the current block into three sub-blocks as shown in (c) and (f) of FIG. 26. The sizes of the three sub-blocks may be different. For example, in the three sub-blocks, the large sub-block may have twice the width / height of the small sub-block. Also, for example, the selected partition mode can divide the current block into two sub-blocks of different sizes (asymmetric binary tree) as shown in (a), (b), (d), and (e) of FIG. 26. For example, if an asymmetric binary tree is used, the large sub-blocks may have three times the width / height of the small sub-blocks.

[0208] FIG. 27 shows an example of dividing a block of NxN pixels using a partition mode selected when the second partition mode is not selected as shown in step (3) in the second embodiment. As shown in FIG. 27, in step (1), a block of Nx2N pixels is divided horizontally into two sub-blocks of NxN pixels, and in step (2), the upper block of NxN pixels is divided horizontally into two sub-blocks of NxN / 2 pixels. In step (3), the selected partition mode for the current block of NxN pixels (the lower block in this example) can be used to divide the current block into three sub-blocks as shown in (c) and (f) of FIG. 27. The sizes of the three sub-blocks may be different. For example, in the three sub-blocks, the large sub-block may have twice the width / height of the small sub-block. Also, for example, the selected partition mode can divide the current block into two sub-blocks of different sizes (asymmetric binary tree) as shown in (a), (b), (d), and (e) of FIG. 27. For example, if an asymmetric binary tree is used, the large sub-blocks may have three times the width / height of the small sub-blocks.

[0209] Figure 17 illustrates possible locations of the first parameter in a compressed video stream. As illustrated in Figure 17, the first parameter may be located in a video parameter set, a sequence parameter set, a picture parameter set, a slice header, or a coding tree unit. The first parameter may indicate how to divide the block into multiple sub-blocks. For example, the first parameter may include a flag indicating whether to divide the block horizontally or vertically. The first parameter may also include a parameter indicating whether to divide the block into two or more sub-blocks.

[0210] FIG. 18 illustrates possible locations of the second parameter in the compressed video stream. As shown in FIG. 18, the second parameter may be located in a video parameter set, a sequence parameter set, a picture parameter set, a slice header, or a coding tree unit. The second parameter may indicate how to divide the block into multiple sub-blocks. For example, the second parameter may include a flag indicating whether to divide the block horizontally or vertically. The second parameter may also include a parameter indicating whether to divide the block into two or more sub-blocks. The second parameter is located following the first parameter in the bitstream, as shown in FIG. 19.

[0211] The first block and the second block are different blocks. The first block and the second block may be included in the same frame. For example, the first block may be an adjacent block above the second block. Also, for example, the first block may be an adjacent block to the left of the second block.

[0212] In step S2006, the second block is divided into sub-blocks using the selected partition mode. In step S2007, the divided blocks are decoded.

[0213] [Decryption device] FIG. 16 is a block diagram showing a structure of a video / picture decoding device according to the second or third embodiment. In FIG.

[0214] The video decoding device 6000 is a device for decoding an input coded bit stream for each block and outputting a video / image. As shown in Fig. 16, the video decoding device 6000 includes an entropy decoding unit 6001, an inverse quantization unit 6002, an inverse transform unit 6003, a block memory 6004, a frame memory 6005, an intra prediction unit 6006, an inter prediction unit 6007, and a block division determination unit 6008.

[0215] An input coded bitstream is input to an entropy decoding unit 6001. After the input coded bitstream is input to the entropy decoding unit 6001, the entropy decoding unit 6001 decodes the input coded bitstream, outputs parameters to a block division determining unit 6008, and outputs a decoded value to an inverse quantization unit 6002.

[0216] The inverse quantization unit 6002 inverse quantizes the decoded value and outputs the frequency coefficient to the inverse transform unit 6003. The inverse transform unit 6003 performs an inverse frequency transform on the frequency coefficient to convert the frequency coefficient into a sample value based on the block partition mode derived by the block partition determination unit 6008, and outputs the sample value to the adder. The block partition mode can be associated with a block partition mode, a block partition type, or a block partition direction. The adder adds the sample value to the predicted video / image value output from the intra / inter prediction unit 6006, 6007, outputs the added value to a display, and outputs the added value to the block memory 6004 or frame memory 6005 for further prediction. The block partition determination unit 6008 collects block information from the block memory 6004 or frame memory 6005, and derives a block partition mode using parameters decoded by the entropy decoding unit 6001. Using the derived block partition mode, the block is divided into multiple sub-blocks. Furthermore, the intra / inter prediction units 6006 and 6007 predict the video / image area of ​​the block to be decoded from the video / image stored in the block memory 6004 or the video / image in the frame memory 6005 reconstructed in the block partition mode derived by the block partition determination unit 6008.

[0217] (Embodiment 3) The encoding process and the decoding process according to the third embodiment will be specifically described with reference to Fig. 13 and Fig. 14. The encoding device and the decoding device according to the third embodiment will be specifically described with reference to Fig. 15 and Fig. 16.

[0218] [Encoding process] FIG. 13 shows a video encoding process according to the third embodiment.

[0219] First, in step S3001, a first parameter that identifies a partition type for dividing a first block into sub-blocks from among a plurality of partition types is written into a bitstream.

[0220] In the next step S3002, a second parameter indicating a partition direction is written into the bitstream. The second parameter is placed following the first parameter in the bitstream. The partition type together with the partition direction may constitute a partition mode. The partition type indicates the number of sub-blocks and the partition ratio for dividing the block.

[0221] FIG. 29 shows an example of partition types and partition directions for dividing a block of NxN pixels in the third embodiment. In FIG. 29, (1), (2), (3), and (4) are different partition types, (1a), (2a), (3a), and (4a) are partition modes with different partition types in the vertical partition direction, and (1b), (2b), (3b), and (4b) are partition modes with different partition types in the horizontal partition direction. As shown in FIG. 29, when the partition ratio is 1:1 and the block is divided into a symmetric binary tree (i.e., two sub-blocks) in the vertical direction, the block of NxN pixels is divided using partition mode (1a). When the partition ratio is 1:1 and the block is divided into a symmetric binary tree (i.e., two sub-blocks) in the horizontal direction, the block of NxN pixels is divided using partition mode (1b). A block of NxN pixels is partitioned using partition mode (2a) if the partition ratio is 1:3 and the block is divided vertically with an asymmetric binary tree (i.e., two sub-blocks). A block of NxN pixels is partitioned using partition mode (2b) if the partition ratio is 1:3 and the block is divided horizontally with an asymmetric binary tree (i.e., two sub-blocks). A block of NxN pixels is partitioned using partition mode (3a) if the partition ratio is 3:1 and the block is divided vertically with an asymmetric binary tree (i.e., two sub-blocks). A block of NxN pixels is partitioned using partition mode (3b) if the partition ratio is 3:1 and the block is divided horizontally with an asymmetric binary tree (i.e., two sub-blocks). A block of NxN pixels is partitioned using partition mode (4a) if the partition ratio is 1:2:1 and the block is divided vertically with a ternary tree (i.e., three sub-blocks). A block of NxN pixels is partitioned using partition mode (4b) where the partition ratio is 1:2:1 and the partition is horizontally ternary tree (ie, 3 sub-blocks).

[0222] Figure 17 illustrates possible locations of the first parameter in a compressed video stream. As illustrated in Figure 17, the first parameter may be located in a video parameter set, a sequence parameter set, a picture parameter set, a slice header, or a coding tree unit. The first parameter may indicate how to divide the block into multiple sub-blocks. For example, the first parameter may include a flag indicating whether to divide the block horizontally or vertically. The first parameter may also include a parameter indicating whether to divide the block into two or more sub-blocks.

[0223] FIG. 18 illustrates possible locations of the second parameter in the compressed video stream. As shown in FIG. 18, the second parameter may be located in a video parameter set, a sequence parameter set, a picture parameter set, a slice header, or a coding tree unit. The second parameter may indicate how to divide the block into multiple sub-blocks. For example, the second parameter may include a flag indicating whether to divide the block horizontally or vertically. The second parameter may also include a parameter indicating whether to divide the block into two or more sub-blocks. The second parameter is located following the first parameter in the bitstream, as shown in FIG. 19.

[0224] FIG. 30 shows the advantage of coding the partition type before the partition direction compared to coding the partition direction before the partition type. In this example, there is no need to code the partition direction when the horizontal partition direction is disabled due to an unsupported size (16x2 pixels). In this example, the partition direction is determined as the vertical partition direction and the horizontal partition direction is disabled. Coding the partition type before the partition direction saves code bits due to coding the partition direction compared to coding the partition direction before the partition type.

[0225] In this way, whether or not a block can be divided in each of the horizontal and vertical directions may be determined based on a predetermined condition of whether or not a block can be divided. Then, if it is determined that a block can be divided in only one of the horizontal and vertical directions, writing in the partition direction to the bit stream may be skipped. Furthermore, if it is determined that a block cannot be divided in both the horizontal and vertical directions, writing in the partition direction and the partition type to the bit stream may be skipped.

[0226] The predetermined condition for whether a block can be divided or not is defined by, for example, a size (number of pixels) or the number of divisions. The condition for whether a block can be divided or not may be predefined in a standard. The condition for whether a block can be divided or not may be included in a video parameter set, a sequence parameter set, a picture parameter set, a slice header, or a coding tree unit. The condition for whether a block can be divided or not may be fixed for all blocks, or may be dynamically switched depending on the characteristics of the block (e.g., luminance and chrominance blocks) or the characteristics of the picture (e.g., I, P, B pictures), etc.

[0227] In step S3003, the block is divided into sub-blocks using the identified partition type and the indicated partition direction. In step S3004, the divided block is coded.

[0228] [Encoding device] FIG. 15 is a block diagram showing a structure of a video / image coding device according to the second or third embodiment. In FIG.

[0229] The video encoding device 5000 is a device for encoding an input video / image for each block to generate an encoded output bitstream. As shown in Fig. 15, the video encoding device 5000 includes a transform unit 5001, a quantization unit 5002, an inverse quantization unit 5003, an inverse transform unit 5004, a block memory 5005, a frame memory 5006, an intra prediction unit 5007, an inter prediction unit 5008, an entropy encoding unit 5009, and a block division determination unit 5010.

[0230] An input image is input to the adder, and an added value is output to the conversion unit 5001. The conversion unit 5001 converts the added value into a frequency coefficient based on the block partition type and direction derived by the block partition determination unit 5010, and outputs the frequency coefficient to the quantization unit 5002. The block partition type and direction may be associated with a block partition mode, a block partition type, or a block partition direction. The quantization unit 5002 quantizes the input quantization coefficient, and outputs the quantization value to the inverse quantization unit 5003 and the entropy coding unit 5009.

[0231] The inverse quantization unit 5003 inversely quantizes the quantized value output from the quantization unit 5002, and outputs the frequency coefficients to the inverse transform unit 5004. The inverse transform unit 5004 performs an inverse frequency transform on the frequency coefficients based on the block partition type and direction derived by the block partition determination unit 5010, converts the frequency coefficients into sample values ​​of a bit stream, and outputs the sample values ​​to an adder.

[0232] The adder adds the sample values ​​of the bitstream output from the inverse transform unit 5004 to the predicted video / image values ​​output from the intra / inter prediction units 5007, 5008, and outputs the sum to the block memory 5005 or the frame memory 5006 for further prediction. The block partition determination unit 5010 collects block information from the block memory 5005 or the frame memory 5006, and derives a block partition type and direction, and parameters related to the block partition type and direction. Using the derived block partition type and direction, the block is divided into multiple sub-blocks. The intra / inter prediction units 5007, 5008 search among the video / image stored in the block memory 5005 or the video / image in the frame memory 5006 reconstructed with the block partition type and direction derived by the block partition determination unit 5010, and estimate, for example, a video / image area that is most similar to the input video / image to be predicted.

[0233] The entropy coding unit 5009 codes the quantized value output from the quantization unit 5002, codes the parameters from the block division determination unit 5010, and outputs a bit stream.

[0234] [Decryption process] FIG. 14 shows a video decoding process according to the third embodiment.

[0235] First, in step S4001, a first parameter that identifies a partition type for dividing a first block into sub-blocks from among a plurality of partition types is read from the bitstream.

[0236] In the next step S4002, a second parameter indicating a partition direction is read from the bitstream. The second parameter follows the first parameter in the bitstream. The partition type together with the partition direction may constitute a partition mode. The partition type indicates the number of sub-blocks and the partition ratio to divide the block.

[0237] FIG. 29 shows an example of partition types and partition directions for dividing a block of NxN pixels in the third embodiment. In FIG. 29, (1), (2), (3), and (4) are different partition types, (1a), (2a), (3a), and (4a) are partition modes with different partition types in the vertical partition direction, and (1b), (2b), (3b), and (4b) are partition modes with different partition types in the horizontal partition direction. As shown in FIG. 29, when the partition ratio is 1:1 and the block is divided into a symmetric binary tree (i.e., two sub-blocks) in the vertical direction, the block of NxN pixels is divided using partition mode (1a). When the partition ratio is 1:1 and the block is divided into a symmetric binary tree (i.e., two sub-blocks) in the horizontal direction, the block of NxN pixels is divided using partition mode (1b). A block of NxN pixels is partitioned using partition mode (2a) if the partition ratio is 1:3 and the block is divided vertically with an asymmetric binary tree (i.e., two sub-blocks). A block of NxN pixels is partitioned using partition mode (2b) if the partition ratio is 1:3 and the block is divided horizontally with an asymmetric binary tree (i.e., two sub-blocks). A block of NxN pixels is partitioned using partition mode (3a) if the partition ratio is 3:1 and the block is divided vertically with an asymmetric binary tree (i.e., two sub-blocks). A block of NxN pixels is partitioned using partition mode (3b) if the partition ratio is 3:1 and the block is divided horizontally with an asymmetric binary tree (i.e., two sub-blocks). A block of NxN pixels is partitioned using partition mode (4a) if the partition ratio is 1:2:1 and the block is divided vertically with a ternary tree (i.e., three sub-blocks). A block of NxN pixels is partitioned using partition mode (4b) where the partition ratio is 1:2:1 and the partition is horizontally ternary tree (ie, 3 sub-blocks).

[0238] FIG. 17 illustrates possible locations of the first parameter in a compressed video stream. As illustrated in FIG. 17, the first parameter may be located in a video parameter set, a sequence parameter set, a picture parameter set, a slice header, or a coding tree unit. The first parameter may indicate how to divide the block into multiple sub-blocks. For example, the first parameter may include an identifier of a partition type as described above. For example, the first parameter may include a flag indicating whether to divide the block horizontally or vertically. The first parameter may also include a parameter indicating whether to divide the block into two or more sub-blocks.

[0239] FIG. 18 illustrates possible locations of the second parameter in the compressed video stream. As shown in FIG. 18, the second parameter may be located in a video parameter set, a sequence parameter set, a picture parameter set, a slice header, or a coding tree unit. The second parameter may indicate how to partition the block into multiple sub-blocks. For example, the second parameter may include a flag indicating whether to partition the block horizontally or vertically. That is, the second parameter may include a parameter indicating a partition direction. The second parameter may also include a parameter indicating whether to partition the block into two or more sub-blocks. The second parameter is located following the first parameter in the bitstream, as shown in FIG. 19.

[0240] FIG. 30 shows the advantage of coding the partition type before the partition direction compared to coding the partition direction before the partition type. In this example, there is no need to code the partition direction when the horizontal partition direction is disabled due to an unsupported size (16x2 pixels). In this example, the partition direction is determined as the vertical partition direction and the horizontal partition direction is disabled. Coding the partition type before the partition direction saves code bits due to coding the partition direction compared to coding the partition direction before the partition type.

[0241] In this way, it may be determined whether or not a block can be divided in each of the horizontal and vertical directions based on a predetermined condition of whether or not a block can be divided. Then, when it is determined that a block can be divided in only one of the horizontal and vertical directions, deciphering from the bit stream in the partition direction may be skipped. Furthermore, when it is determined that a block cannot be divided in both the horizontal and vertical directions, deciphering from the bit stream of the partition type in addition to the partition direction may be skipped.

[0242] The predetermined condition for whether a block can be divided or not is defined by, for example, a size (number of pixels) or the number of divisions. The condition for whether a block can be divided or not may be predefined in a standard. The condition for whether a block can be divided or not may be included in a video parameter set, a sequence parameter set, a picture parameter set, a slice header, or a coding tree unit. The condition for whether a block can be divided or not may be fixed for all blocks, or may be dynamically switched depending on the characteristics of the block (e.g., luminance and chrominance blocks) or the characteristics of the picture (e.g., I, P, B pictures), etc.

[0243] In step S4003, the block is divided into sub-blocks using the identified partition type and the indicated partition direction. In step S4004, the divided block is decoded.

[0244] [Decryption device] FIG. 16 is a block diagram showing a structure of a video / picture decoding device according to the second or third embodiment. In FIG.

[0245] The video decoding device 6000 is a device for decoding an input coded bit stream for each block and outputting a video / image. As shown in Fig. 16, the video decoding device 6000 includes an entropy decoding unit 6001, an inverse quantization unit 6002, an inverse transform unit 6003, a block memory 6004, a frame memory 6005, an intra prediction unit 6006, an inter prediction unit 6007, and a block division determination unit 6008.

[0246] An input coded bitstream is input to an entropy decoding unit 6001. After the input coded bitstream is input to the entropy decoding unit 6001, the entropy decoding unit 6001 decodes the input coded bitstream, outputs parameters to a block division determining unit 6008, and outputs a decoded value to an inverse quantization unit 6002.

[0247] The inverse quantization unit 6002 inverse quantizes the decoded value and outputs the frequency coefficient to the inverse transform unit 6003. The inverse transform unit 6003 performs an inverse frequency transform on the frequency coefficient to convert the frequency coefficient into a sample value based on the block partition type and direction derived by the block partition determination unit 6008, and outputs the sample value to the adder. The block partition type and direction can be associated with a block partition mode, a block partition type, or a block partition direction. The adder adds the sample value to the predicted video / image value output from the intra / inter prediction unit 6006, 6007, outputs the added value to the display, and outputs the added value to the block memory 6004 or the frame memory 6005 for further prediction. The block partition determination unit 6008 collects block information from the block memory 6004 or the frame memory 6005, and derives a block partition type and direction using the parameters decoded by the entropy decoding unit 6001. Using the derived block partition type and direction, the block is divided into multiple sub-blocks. Furthermore, the intra / inter prediction units 6006 and 6007 predict the video / image area of ​​the block to be decoded from the video / image stored in the block memory 6004 or the video / image in the frame memory 6005 reconstructed using the block partition type and direction derived by the block partition determination unit 6008.

[0248] (Embodiment 4) In each of the above embodiments, each of the functional blocks can usually be realized by an MPU, a memory, etc. Furthermore, the processing by each of the functional blocks is usually realized by a program execution unit such as a processor reading and executing software (programs) recorded on a recording medium such as a ROM. The software may be distributed by downloading, etc., or may be recorded on a recording medium such as a semiconductor memory and distributed. Of course, each functional block can also be realized by hardware (dedicated circuitry).

[0249] Furthermore, the processes described in each embodiment may be realized by centralized processing using a single device (system), or may be realized by distributed processing using multiple devices. The processor that executes the above program may be either single or multiple. That is, centralized processing or distributed processing may be performed.

[0250] The aspects of the present disclosure are not limited to the above-described examples, and various modifications are possible, which are also included within the scope of the aspects of the present disclosure.

[0251] Further, here, application examples of the video coding method (image coding method) or video decoding method (image decoding method) shown in each of the above embodiments and a system using the same will be described. The system is characterized by having an image coding device using the image coding method, an image decoding device using the image decoding method, and an image coding / decoding device equipped with both. Other configurations in the system can be appropriately changed depending on the case.

[0252] [Usage example] 33 is a diagram showing the overall configuration of a content supply system ex100 that realizes a content distribution service. The area where communication services are provided is divided into cells of a desired size, and base stations ex106, ex107, ex108, ex109, and ex110, which are fixed wireless stations, are installed in each cell.

[0253] In this content supply system ex100, devices such as a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, and a smartphone ex115 are connected to the Internet ex101 via an Internet service provider ex102 or a communication network ex104, and base stations ex106 to ex110. The content supply system ex100 may be configured to connect any of the above elements in combination. Each device may be directly or indirectly connected to each other via a telephone network or short-distance wireless communication, without going through the base stations ex106 to ex110, which are fixed wireless stations. In addition, the streaming server ex103 is connected to each device such as a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, and a smartphone ex115 via the Internet ex101, etc. In addition, the streaming server ex103 is connected to a terminal in a hot spot in an airplane ex117, etc., via a satellite ex116.

[0254] Instead of the base stations ex106 to ex110, wireless access points or hot spots may be used. The streaming server ex103 may be directly connected to the communication network ex104 without going through the Internet ex101 or the Internet service provider ex102, or may be directly connected to an airplane ex117 without going through a satellite ex116.

[0255] The camera ex113 is a device capable of taking still images and videos, such as a digital camera. The smartphone ex115 is a smartphone, a mobile phone, or a PHS (Personal Handyphone System) that is compatible with a mobile communication system generally called 2G, 3G, 3.9G, 4G, and 5G in the future.

[0256] The home appliance ex118 is a refrigerator or an appliance included in a home fuel cell cogeneration system.

[0257] In the content supply system ex100, a terminal having a photographing function is connected to a streaming server ex103 via a base station ex106 or the like, thereby enabling live distribution and the like. In live distribution, a terminal (such as a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, a smartphone ex115, and a terminal in an airplane ex117) performs the encoding process described in each of the above embodiments on still image or video content photographed by a user using the terminal, multiplexes the video data obtained by the encoding with audio data obtained by encoding audio corresponding to the video, and transmits the obtained data to the streaming server ex103. That is, each terminal functions as an image encoding device according to one aspect of the present disclosure.

[0258] Meanwhile, the streaming server ex103 streams the transmitted content data to the requesting client. The client is a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, a smartphone ex115, a terminal in an airplane ex117, or the like, capable of decoding the encoded data. Each device that receives the distributed data decodes and plays back the received data. That is, each device functions as an image decoding device according to one aspect of the present disclosure.

[0259] [Distributed processing] The streaming server ex103 may be a plurality of servers or computers that process, record, and distribute data in a distributed manner. For example, the streaming server ex103 may be realized by a CDN (Contents Delivery Network), and content distribution may be realized by a network that connects a large number of edge servers distributed around the world. In a CDN, an edge server that is physically close to the client is dynamically assigned according to the client. The content is cached and distributed to the edge server, thereby reducing delays. In addition, when an error occurs or the communication state changes due to an increase in traffic, the processing can be distributed among multiple edge servers, the distribution entity can be switched to another edge server, or distribution can be continued by bypassing the part of the network where a failure has occurred, thereby realizing high-speed and stable distribution.

[0260] In addition to the distributed processing of the distribution itself, the encoding processing of the captured data may be performed by each terminal, may be performed by the server side, or may be shared among the terminals. As an example, in the encoding processing, a processing loop is generally performed twice. In the first loop, the complexity of the image or the amount of code is detected for each frame or scene. In the second loop, processing is performed to maintain the image quality and improve the encoding efficiency. For example, the terminal performs the first encoding processing, and the server side that receives the content performs the second encoding processing, thereby improving the quality and efficiency of the content while reducing the processing load on each terminal. In this case, if there is a request to receive and decode almost in real time, the data encoded once by the terminal can be received and played back by other terminals, making it possible to perform more flexible real-time distribution.

[0261] As another example, the camera ex113 etc. extracts features from an image, compresses data related to the features as metadata, and transmits the compressed data to the server. The server performs compression according to the meaning of the image, for example, by determining the importance of an object from the features and switching the quantization precision. The feature data is particularly effective in improving the precision and efficiency of motion vector prediction when the server performs recompression. Alternatively, the terminal may perform simple encoding such as VLC (variable length coding), and the server may perform encoding with a large processing load such as CABAC (context-adaptive binary arithmetic coding).

[0262] As another example, in a stadium, a shopping mall, a factory, etc., there may be a plurality of video data in which almost the same scene has been shot by a plurality of terminals. In this case, using the plurality of terminals that shot the video and, as necessary, other terminals and servers that did not shoot the video, coding processing is assigned to each of them, for example, in units of GOPs (Group of Pictures), in units of pictures, or in units of tiles obtained by dividing a picture, for distributed processing. This reduces delays and realizes better real-time performance.

[0263] In addition, since the multiple video data are of almost the same scene, the server may manage and / or instruct the video data shot by each terminal to be mutually referenced. Alternatively, the server may receive the encoded data from each terminal and change the reference relationship between the multiple data, or correct or replace the pictures themselves and re-encode them. This makes it possible to generate a stream with improved quality and efficiency for each piece of data.

[0264] The server may also perform transcoding to change the encoding format of the video data before distributing it. For example, the server may convert an MPEG-based encoding format into a VP-based encoding format, or convert H.264 into H.265.

[0265] In this way, the encoding process can be performed by a terminal or one or more servers. Therefore, in the following, descriptions such as "server" or "terminal" are used to indicate the entity performing the processing, but some or all of the processing performed by the server may be performed by the terminal, and some or all of the processing performed by the terminal may be performed by the server. The same applies to the decoding process.

[0266] [3D, multi-angle] In recent years, it has become common to integrate and use images or videos of different scenes or the same scene taken from different angles by multiple devices such as cameras ex113 and / or smartphones ex115 that are almost synchronized with each other. The videos taken by the devices are integrated based on the relative positional relationship between the devices that is obtained separately, or on areas where feature points included in the videos match.

[0267] The server may not only encode 2D video, but also encode still images automatically or at a time specified by the user based on scene analysis of the video and transmit them to the receiving terminal. If the server can obtain the relative positional relationship between the shooting terminals, the server may generate a 3D shape of the scene based on not only 2D video but also images of the same scene captured from different angles. The server may separately encode 3D data generated by point clouds, etc., or may generate images to be transmitted to the receiving terminal by selecting or reconstructing images from images captured by multiple terminals based on the results of recognizing or tracking people or objects using the 3D data.

[0268] In this way, the user can enjoy a scene by selecting any video corresponding to each shooting terminal, or can enjoy content in which a video from any viewpoint is cut out from 3D data reconstructed using multiple images or videos. Furthermore, sound may be collected from multiple different angles, just like the video, and the server may multiplex the sound from a specific angle or space with the video and transmit it in accordance with the video.

[0269] In recent years, content that associates the real world with a virtual world, such as Virtual Reality (VR) and Augmented Reality (AR), has also become popular. In the case of VR images, the server creates viewpoint images for the right eye and the left eye, respectively, and may perform encoding that allows reference between each viewpoint video using Multi-View Coding (MVC) or the like, or may encode them as separate streams without mutual reference. When decoding the separate streams, it is preferable to play them in synchronization with each other so that a virtual three-dimensional space is reproduced according to the user's viewpoint.

[0270] In the case of an AR image, the server superimposes virtual object information in the virtual space on camera information in the real space based on the three-dimensional position or the movement of the user's viewpoint. The decoding device may obtain or hold virtual object information and three-dimensional data, generate a two-dimensional image according to the movement of the user's viewpoint, and smoothly connect them to create superimposed data. Alternatively, the decoding device may transmit the movement of the user's viewpoint to the server in addition to a request for virtual object information, and the server may create superimposed data according to the movement of the viewpoint received from the three-dimensional data held by the server, encode the superimposed data, and deliver it to the decoding device. Note that the superimposed data has an α value indicating the transparency in addition to RGB, and the server may set the α value of the part other than the object created from the three-dimensional data to 0, etc., and encode the data in a state in which the part is transparent. Alternatively, the server may generate data in which a predetermined value of RGB value is set to the background like a chromakey, and the part other than the object is the background color.

[0271] Similarly, the decoding process of the distributed data may be performed by each client terminal, or may be performed by the server side, or may be shared among them. As an example, a certain terminal may once send a reception request to the server, and the content corresponding to the request may be received by other terminals, decoded, and the decoded signal may be transmitted to a device having a display. By distributing the processing and selecting appropriate content regardless of the performance of the communication-enabled terminals themselves, data with good image quality can be reproduced. In another example, while large-sized image data is received by a TV or the like, a part of the area, such as tiles into which the picture is divided, may be decoded and displayed on the viewer's personal terminal. This allows the viewer to share the overall picture while checking his / her own area of ​​responsibility or the area he / she wants to check in more detail at hand.

[0272] In the future, it is expected that content will be seamlessly received by switching appropriate data for the currently connected communication using delivery system standards such as MPEG-DASH under circumstances where multiple short-distance, medium-distance, or long-distance wireless communication is available, regardless of whether indoors or outdoors. This allows users to freely select and switch in real time not only their own terminals but also decoding devices or display devices such as displays installed indoors and outdoors. In addition, decoding can be performed while switching the decoding device and the display device based on the user's location information, etc. This makes it possible to move while displaying map information on the wall or part of the ground of a neighboring building where a displayable device is embedded while moving to a destination. It is also possible to switch the bit rate of the received data based on the accessibility of the encoded data on the network, such as when the encoded data is cached on a server that can be accessed from the receiving terminal in a short time, or when it is copied to an edge server in a content delivery service.

[0273] [Scalable Coding] The switching of contents will be described using a scalable stream compressed and coded by applying the video coding method shown in each of the above embodiments, as shown in FIG. 34. The server may have multiple streams with the same content but different qualities as individual streams, but may be configured to switch contents by taking advantage of the characteristics of a temporal / spatial scalable stream realized by coding in layers as shown in the figure. In other words, the decoding side can freely switch between low-resolution content and high-resolution content by determining which layer to decode according to an internal factor such as performance and an external factor such as the state of the communication band. For example, if you want to continue watching a video you were watching on your smartphone ex115 while on the move on a device such as an Internet TV after you get home, the device can decode the same stream up to a different layer, reducing the burden on the server side.

[0274] Furthermore, as described above, in addition to the configuration that realizes scalability in which pictures are coded for each layer and an enhancement layer exists above a base layer, the enhancement layer may include meta-information based on image statistics, etc., and the decoding side may generate high-quality content by super-resolving pictures of the base layer based on the meta-information. Super-resolution may be either an improvement in the signal-to-noise ratio at the same resolution or an increase in resolution. The meta-information includes information for specifying linear or nonlinear filter coefficients used in the super-resolution process, or information for specifying parameter values ​​in the filter process, machine learning, or least squares calculation used in the super-resolution process.

[0275] Alternatively, a picture may be divided into tiles or the like according to the meaning of an object in an image, and the decoding side may select a tile to decode to decode only a part of the area. Also, by storing the attribute of an object (person, car, ball, etc.) and its position in a video (coordinate position in the same image, etc.) as meta information, the decoding side can identify the position of a desired object based on the meta information and determine the tile containing the object. For example, as shown in FIG. 35, the meta information is stored using a data storage structure different from pixel data such as an SEI message in HEVC. This meta information indicates, for example, the position, size, or color of a main object.

[0276] Meta information may also be stored in units consisting of multiple pictures, such as streams, sequences, or random access units, etc. This allows the decoding side to obtain the time when a specific person appears in the video, and by combining this with picture-by-picture information, it is possible to identify the picture in which an object exists and the position of the object within the picture.

[0277] [Web page optimization] FIG. 36 is a diagram showing an example of a display screen of a web page in a computer ex111 or the like. FIG. 37 is a diagram showing an example of a display screen of a web page in a smartphone ex115 or the like. As shown in FIG. 36 and FIG. 37, a web page may include multiple link images that are links to image content, and the appearance of the web page differs depending on the device used to view the page. When multiple link images are visible on the screen, the display device (decoding device) displays a still image or I-picture that each content has as a link image, displays a video such as a gif animation using multiple still images or I-pictures, or receives only the base layer to decode and display the video, until the user explicitly selects the link image, or until the link image approaches the center of the screen or the entire link image enters the screen.

[0278] When a link image is selected by a user, the display device gives top priority to decoding the base layer. If the HTML constituting the web page contains information indicating that the content is scalable, the display device may decode up to the enhancement layer. In order to ensure real-time performance, before selection or when the communication bandwidth is very tight, the display device decodes and displays only forward-reference pictures (I-pictures, P-pictures, and B-pictures with forward reference only), thereby reducing the delay between the decoding time of the first picture and the display time (the delay from the start of decoding the content to the start of display). The display device may also ignore the reference relationship between pictures and roughly decode all B-pictures and P-pictures with forward reference, and perform normal decoding as the number of received pictures increases over time.

[0279] [Automatic driving] Furthermore, when transmitting and receiving still image or video data such as 2D or 3D map information for automatic driving or driving assistance of a vehicle, the receiving terminal may receive weather or construction information as meta information in addition to image data belonging to one or more layers, and may associate and decode these. Note that the meta information may belong to a layer, or may simply be multiplexed with the image data.

[0280] In this case, since a car, drone, or airplane including a receiving terminal moves, the receiving terminal can realize seamless reception and decoding while switching between base stations ex106 to ex110 by transmitting the location information of the receiving terminal at the time of a reception request. Also, the receiving terminal can dynamically switch how much meta information to receive or how much to update map information according to the user's selection, the user's situation, or the state of the communication band.

[0281] In this manner, in the content supply system ex100, the client can receive, decode, and play back encoded information transmitted by the user in real time.

[0282] [Distribution of personal content] Furthermore, the content supply system ex100 allows not only high-quality, long-duration content from video distributors, but also low-quality, short-duration content from individuals via unicast or multicast distribution. It is expected that such personal content will continue to increase in the future. To improve the quality of personal content, the server may perform editing processing before encoding processing. This can be achieved, for example, by the following configuration.

[0283] During shooting, in real time or after accumulating, the server performs recognition processing such as shooting errors, scene search, semantic analysis, and object detection from the original image or encoded data. Then, based on the recognition result, the server manually or automatically corrects out-of-focus or camera shake, deletes less important scenes such as scenes that are less bright than other pictures or are out of focus, emphasizes object edges, changes color, and performs other editing. The server encodes the edited data based on the editing result. It is also known that if the shooting time is too long, the viewer rating will decrease, and the server may automatically clip not only scenes of less importance as described above but also scenes with little movement based on the image processing result so that the content will be within a specific time range depending on the shooting time. Alternatively, the server may generate a digest based on the result of the semantic analysis of the scene and encode it.

[0284] In addition, there are cases where personal contents contain images that infringe copyrights, moral rights, portrait rights, etc., and the scope of sharing may exceed the intended scope, which may be inconvenient for individuals. Therefore, for example, the server may change the image to an unfocused image of a person's face on the periphery of the screen, or the inside of a house, and encode it. The server may also recognize whether the image to be encoded contains a face of a person other than a person registered in advance, and if so, may perform processing such as blurring the face. Alternatively, as pre-processing or post-processing of encoding, the user may specify a person or background area that he or she wishes to process in the image from the viewpoint of copyright, etc., and the server may replace the specified area with another image or blur the focus. If it is a person, the image of the face part can be replaced while tracking the person in the video.

[0285] In addition, since viewing of personal content with a small amount of data requires real-time performance, the decoding device first receives the base layer as a top priority, and performs decoding and playback, although this depends on the bandwidth. The decoding device may receive an enhancement layer during this time, and when playback is looped or otherwise played two or more times, play high-quality video including the enhancement layer. In this way, if the stream is scalably encoded, it is possible to provide an experience in which the video is rough when not selected or when viewing begins, but the stream gradually becomes smarter and the image quality improves. In addition to scalable encoding, a similar experience can be provided even if a rough stream played the first time and a second stream encoded with reference to the first video are configured as one stream.

[0286] [Other use cases] Moreover, these encoding or decoding processes are generally processed in an LSIex500 possessed by each terminal. The LSIex500 may be a single chip or may be configured with multiple chips. Note that software for encoding or decoding moving images may be incorporated into some kind of recording medium (such as a CD-ROM, a flexible disk, or a hard disk) that can be read by the computer ex111 or the like, and the encoding or decoding process may be performed using the software. Furthermore, if the smartphone ex115 has a camera, video data captured by the camera may be transmitted. The video data in this case is data that has been encoded by the LSIex500 possessed by the smartphone ex115.

[0287] The LSIex500 may be configured to download and activate application software. In this case, the terminal first determines whether the terminal supports the encoding method of the content or has the ability to execute a specific service. If the terminal does not support the encoding method of the content or does not have the ability to execute a specific service, the terminal downloads a codec or application software, and then acquires and plays the content.

[0288] Furthermore, at least one of the video encoding device (image encoding device) or video decoding device (image decoding device) of each of the above-mentioned embodiments can be incorporated into a digital broadcasting system, not limited to the content supply system ex100 via the Internet ex101. Since multiplexed data in which video and audio are multiplexed is carried and transmitted on broadcasting radio waves using a satellite or the like, there is a difference in that it is more suitable for multicast compared to the content supply system ex100, which has a configuration that is easy to use for unicast, but similar applications are possible with regard to the encoding process and the decoding process.

[0289] [Hardware configuration] FIG. 38 is a diagram showing a smartphone ex115. FIG. 39 is a diagram showing a configuration example of the smartphone ex115. The smartphone ex115 includes an antenna ex450 for transmitting and receiving radio waves to and from the base station ex110, a camera unit ex465 capable of taking videos and still images, and a display unit ex458 for displaying the video captured by the camera unit ex465 and the decoded data of the video received by the antenna ex450. The smartphone ex115 further includes an operation unit ex466 such as a touch panel, an audio output unit ex457 such as a speaker for outputting audio or sound, an audio input unit ex456 such as a microphone for inputting audio, a memory unit ex467 capable of storing encoded data such as captured video or still images, recorded audio, received video or still images, and e-mail, or decoded data, and a slot unit ex464 which is an interface unit with a SIMex468 for identifying a user and authenticating access to various data including a network. In addition, an external memory may be used instead of the memory unit ex467.

[0290] In addition, a main control unit ex460, which comprehensively controls the display unit ex458 and the operation unit ex466, etc., is connected to a power supply circuit unit ex461, an operation input control unit ex462, a video signal processing unit ex455, a camera interface unit ex463, a display control unit ex459, a modulation / demodulation unit ex452, a multiplexing / separation unit ex453, an audio signal processing unit ex454, a slot unit ex464, and a memory unit ex467 via a bus ex470.

[0291] When the power key is turned on by a user's operation, the power supply circuit unit ex461 starts up the smartphone ex115 into an operational state by supplying power to each unit from the battery pack.

[0292] The smartphone ex115 performs processes such as telephone calls and data communications under the control of a main control unit ex460 having a CPU, a ROM, and a RAM. During a telephone call, a voice signal collected by a voice input unit ex456 is converted into a digital voice signal by a voice signal processing unit ex454, which is then subjected to spectrum spreading processing by a modulation / demodulation unit ex452, and the digital-to-analog conversion processing and frequency conversion processing by a transmission / reception unit ex451 is then transmitted via an antenna ex450. In addition, the received data is amplified, and subjected to frequency conversion processing and analog-to-digital conversion processing, and the spectrum inverse spreading processing by a modulation / demodulation unit ex452 is then performed, and the analog voice signal is converted into an analog voice signal by a voice signal processing unit ex454, which is then output from a voice output unit ex457. During a data communication mode, text, still images, or video data is sent to the main control unit ex460 via an operation input control unit ex462 by operating an operation unit ex466 or the like of the main unit, and transmission and reception processing is performed in the same manner. When transmitting video, still images, or video and audio in the data communication mode, the video signal processing unit ex455 compresses and encodes the video signal stored in the memory unit ex467 or the video signal input from the camera unit ex465 by the moving image encoding method shown in each of the above embodiments, and sends the encoded video data to the multiplexing / separation unit ex453. The audio signal processing unit ex454 also encodes the audio signal collected by the audio input unit ex456 while the camera unit ex465 is capturing the video or still images, and sends the encoded audio data to the multiplexing / separation unit ex453. The multiplexing / separation unit ex453 multiplexes the encoded video data and the encoded audio data by a predetermined method, and performs modulation and conversion processing in the modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmission / reception unit ex451, and transmits the data via the antenna ex450.

[0293] When receiving a video attached to an e-mail or chat, or a video linked to a web page, etc., in order to decode the multiplexed data received via the antenna ex450, the multiplexing / separation unit ex453 separates the multiplexed data into a bit stream of video data and a bit stream of audio data by separating the multiplexed data, and supplies the encoded video data to the video signal processing unit ex455 via the synchronization bus ex470, and supplies the encoded audio data to the audio signal processing unit ex454. The video signal processing unit ex455 decodes the video signal by a video decoding method corresponding to the video encoding method shown in each of the above embodiments, and displays the video or still image contained in the linked video file on the display unit ex458 via the display control unit ex459. The audio signal processing unit ex454 also decodes the audio signal, and audio is output from the audio output unit ex457. Note that since real-time streaming is widespread, there may be cases where audio playback is socially inappropriate depending on the user's situation. Therefore, as an initial value, a configuration in which only video data is played without playing audio signals is preferable. The audio may be played in sync only when the user performs an operation such as clicking on the video data.

[0294] In addition, although the smartphone ex115 has been described as an example here, three types of implementation formats are possible for the terminal: a transmitting / receiving terminal having both an encoder and a decoder, a transmitting terminal having only an encoder, and a receiving terminal having only a decoder. Furthermore, in the digital broadcasting system, multiplexed data in which audio data and the like are multiplexed with video data is received or transmitted, but the multiplexed data may include text data related to the video in addition to audio data, or the video data itself may be received or transmitted instead of the multiplexed data.

[0295] Although the main control unit ex460 including the CPU controls the encoding or decoding process, terminals often have a GPU. Therefore, a configuration may be used in which a wide area is processed collectively by utilizing the performance of the GPU using a memory shared by the CPU and GPU, or a memory whose addresses are managed so that they can be used in common. This can shorten the encoding time, ensure real-time performance, and achieve low latency. In particular, it is efficient to perform the processing of motion search, deblocking filter, SAO (Sample Adaptive Offset), and transformation and quantization collectively in units such as pictures by the GPU, rather than by the CPU.

[0296] An encoding device of an embodiment of the present disclosure is an encoding device that encodes a picture, and includes a processor and a memory. The processor has a block partition determination unit that divides the picture read from the memory into a plurality of blocks using a block partition mode set that combines one or a plurality of block partition modes that define a partition type, and an encoding unit that encodes the plurality of blocks. The block partition mode set consists of a first block partition mode that defines a partition direction and a partition number for dividing a first block, and a second block partition mode that defines a partition direction and a partition number for dividing a second block, which is one of the blocks obtained after dividing the first block. The block partition determination unit may determine that the second block partition mode only includes a block partition mode with the partition number of 3 when the partition number of the first block partition mode is 3, the second block is a central block among the blocks obtained after dividing the first block, and the partition direction of the second block partition mode is the same as the partition direction of the first block partition mode.

[0297] The parameters for identifying the second block division mode in the encoding device of an embodiment of the present disclosure may include a first flag indicating whether the block is to be divided horizontally or vertically, and may not include a second flag indicating the number of divisions into which the block is to be divided.

[0298] An encoding device of an embodiment of the present disclosure is an encoding device that encodes a picture, and includes a processor and a memory. The processor has a block partition determination unit that divides the picture read from the memory into a plurality of blocks using a block partition mode set that combines one or a plurality of block partition modes that define a partition type, and an encoding unit that encodes the plurality of blocks, and the block partition mode set consists of a first block partition mode that defines a partition direction and a partition number for dividing a first block, and a second block partition mode that defines a partition direction and a partition number for dividing a second block, which is one of the blocks obtained after dividing the first block, and the block partition determination unit may not use the second block partition mode with the partition number of 2 when the partition number of the first block partition mode is 3, the second block is a central block among the blocks obtained after dividing the first block, and the partition direction of the second block partition mode is the same as the partition direction of the first block partition mode.

[0299] An encoding device of an embodiment of the present disclosure is an encoding device that encodes a picture, and includes a processor and a memory, wherein the processor has a block partition determination unit that divides the picture read from the memory into a plurality of blocks using a block partition mode set that combines one or a plurality of block partition modes that define a partition type, and an encoding unit that encodes the plurality of blocks, wherein the block partition mode set includes a first block partition mode and a second block partition mode that each define a partition direction and a partition number, and the block partition determination unit may restrict the use of the second block partition mode in which the partition number is 2.

[0300] The parameters for identifying the second block division mode in the encoding device of an embodiment of the present disclosure may include a first flag indicating whether the block is to be divided horizontally or vertically, and a second flag indicating whether the block is to be divided into two or more parts.

[0301] In the encoding device according to the embodiment of the present disclosure, the parameters may be placed in slice data.

[0302] An encoding device of an embodiment of the present disclosure is an encoding device that encodes a picture, and is equipped with a processor and a memory, wherein the processor has a block partition determination unit that divides the picture read from the memory into a block set consisting of a plurality of blocks using a block partition mode set that combines one or a plurality of block partition modes that define a partition type, and an encoding unit that encodes the plurality of blocks, and when a first block set obtained using a first block partition mode set and a second block set obtained using a second block partition mode set are identical, the block partition determination unit may perform division using only either the first block partition mode set or the second block partition mode set.

[0303] The block partitioning determination unit in the encoding device of an embodiment of the present disclosure may perform partitioning using the block partitioning mode set having the smaller of the first encoding amount and the second encoding amount, based on a first encoding amount of the first block partitioning mode set and a second encoding amount of the second block partitioning mode set.

[0304] The block partitioning determination unit in the encoding device of an embodiment of the present disclosure may, based on a first code amount of the first block partitioning mode set and a second code amount of the second block partitioning mode set, perform partitioning using a block partitioning mode set that appears first in a predetermined order among the first block partitioning mode set and the second block partitioning mode set when the first code amount and the second code amount are equal.

[0305] A decoding device of an embodiment of the present disclosure is a decoding device that decodes an encoded signal, and includes a processor and a memory. The processor has a block partition determination unit that divides the encoded signal read from the memory into a plurality of blocks using a block partition mode set that combines one or a plurality of block partition modes that define a partition type, and a decoding unit that decodes the plurality of blocks. The block partition mode set consists of a first block partition mode that defines a partition direction and a partition number for dividing a first block, and a second block partition mode that defines a partition direction and a partition number for dividing a second block, which is one of the blocks obtained after dividing the first block. The block partition determination unit may determine that the second block partition mode only includes a block partition mode with the partition number of 3 when the partition number of the first block partition mode is 3, the second block is a central block among the blocks obtained after dividing the first block, and the partition direction of the second block partition mode is the same as the partition direction of the first block partition mode.

[0306] The parameters for identifying the second block division mode in the decoding device of an embodiment of the present disclosure may include a first flag indicating whether the block is to be divided horizontally or vertically, and may not include a second flag indicating the number of divisions into which the block is to be divided.

[0307] A decoding device of an embodiment of the present disclosure is a decoding device that decodes an encoded signal, and includes a processor and a memory. The processor has a block partition determination unit that divides the encoded signal read from the memory into a plurality of blocks using a block partition mode set that combines one or a plurality of block partition modes that define a partition type, and a decoding unit that decodes the plurality of blocks. The block partition mode set consists of a first block partition mode that defines a partition direction and a partition number for dividing a first block, and a second block partition mode that defines a partition direction and a partition number for dividing a second block, which is one of the blocks obtained after dividing the first block. The block partition determination unit may not use the second block partition mode with the partition number of 2 when the partition number of the first block partition mode is 3, the second block is a central block among the blocks obtained after dividing the first block, and the partition direction of the second block partition mode is the same as the partition direction of the first block partition mode.

[0308] A decoding device of an embodiment of the present disclosure is a decoding device that decodes an encoded signal, and includes a processor and a memory, wherein the processor has a block partitioning determination unit that divides the encoded signal read from the memory into a plurality of blocks using a block partitioning mode set that combines one or a plurality of block partitioning modes that define a partition type, and a decoding unit that decodes the plurality of blocks, wherein the block partitioning mode set includes a first block partitioning mode and a second block partitioning mode that each define a partitioning direction and a partition number, and the block partitioning determination unit may restrict the use of the second block partitioning mode in which the partition number is 2.

[0309] The parameters for identifying the second block division mode in the decoding device of an embodiment of the present disclosure may include a first flag indicating whether the block is to be divided horizontally or vertically, and a second flag indicating whether the block is to be divided into two or more parts.

[0310] In the decoding device of the embodiment of the present disclosure, the parameters may be placed in slice data.

[0311] A decoding device of an embodiment of the present disclosure is a decoding device that decodes an encoded signal, and is equipped with a processor and a memory, wherein the processor has a block partition determination unit that divides the encoded signal read from the memory into a block set consisting of a plurality of blocks using a block partition mode set that combines one or a plurality of block partition modes that define a partition type, and a decoding unit that decodes the plurality of blocks, and when a first block set obtained using a first block partition mode set and a second block set obtained using a second block partition mode set are identical, the block partition determination unit may perform partitioning using only either the first block partition mode set or the second block partition mode set.

[0312] The block partitioning determination unit in the decoding device of an embodiment of the present disclosure may perform partitioning using the block partitioning mode set having the smaller of the first code amount and the second code amount, based on a first code amount of the first block partitioning mode set and a second code amount of the second block partitioning mode set.

[0313] The block partitioning determination unit in the decoding device of an embodiment of the present disclosure may, based on a first code amount of the first block partitioning mode set and a second code amount of the second block partitioning mode set, perform partitioning using a block partitioning mode set that appears first among the first block partitioning mode set and the second block partitioning mode set in a predetermined order when the first code amount and the second code amount are equal.

[0314] An encoding method of an embodiment of the present disclosure divides a picture read from a memory into a plurality of blocks using a block partition mode set that combines one or a plurality of block partition modes that define a partition type, and encodes the plurality of blocks, wherein the block partition mode set includes a first block partition mode that defines a partition direction and a partition number for dividing a first block, and a second block partition mode that defines a partition direction and a partition number for dividing a second block that is one of the blocks obtained after dividing the first block, and in the division, if the partition number of the first block partition mode is 3, the second block is a central block among the blocks obtained after dividing the first block, and the partition direction of the second block partition mode is the same as the partition direction of the first block partition mode, the second block partition mode may only include a block partition mode with the partition number of 3.

[0315] The parameters for identifying the second block division mode in the encoding method of an embodiment of the present disclosure may include a first flag indicating whether the block is to be divided horizontally or vertically, and may not include a second flag indicating the number of divisions into which the block is to be divided.

[0316] An encoding method of an embodiment of the present disclosure includes a step of dividing a picture read from a memory into a plurality of blocks using a block partition mode set that combines one or a plurality of block partition modes that define a partition type, and a step of encoding the plurality of blocks, wherein the block partition mode set includes a first block partition mode that defines a partition direction and a partition number for dividing a first block, and a second block partition mode that defines a partition direction and a partition number for dividing a second block that is one of the blocks obtained after dividing the first block, and the dividing step does not need to use the second block partition mode with the partition number of 2 if the partition number of the first block partition mode is 3, the second block is a central block among the blocks obtained after dividing the first block, and the partition direction of the second block partition mode is the same as the partition direction of the first block partition mode.

[0317] An encoding method of an embodiment of the present disclosure includes a step of dividing a picture read from a memory into a plurality of blocks using a block partitioning mode set that includes one or a combination of a plurality of block partitioning modes that define a partition type, and a step of encoding the plurality of blocks, wherein the block partitioning mode set includes a first block partitioning mode and a second block partitioning mode that each define a partitioning direction and a partitioning number, and the partitioning step may be restricted to use of the second block partitioning mode in which the partitioning number is 2.

[0318] The parameters for identifying the second block division mode in the encoding method of an embodiment of the present disclosure may include a first flag indicating whether the block is to be divided horizontally or vertically, and a second flag indicating whether the block is to be divided into two or more parts.

[0319] The parameters in the encoding method of the present disclosure may be located in slice data.

[0320] An encoding method of an embodiment of the present disclosure includes a step of dividing a picture read from a memory into a block set consisting of a plurality of blocks using a block partition mode set that is a combination of one or a plurality of block partition modes that define a partition type, and a step of encoding the plurality of blocks, wherein the dividing step may include, when a first block set obtained using a first block partition mode set and a second block set obtained using a second block partition mode set are identical, dividing the picture using only either the first block partition mode set or the second block partition mode set.

[0321] The splitting step in the encoding method of an embodiment of the present disclosure may involve splitting using a block partitioning mode set which has a smaller first code amount or a smaller second code amount, based on a first code amount of the first block partitioning mode set and a second code amount of the second block partitioning mode set.

[0322] The splitting step in the encoding method of an embodiment of the present disclosure may be based on a first code amount of the first block partition mode set and a second code amount of the second block partition mode set, and when the first code amount and the second code amount are equal, splitting may be performed using a block partition mode set that appears first among the first block partition mode set and the second block partition mode set in a predetermined order.

[0323] A decoding method of an embodiment of the present disclosure divides an encoded signal read from a memory into a plurality of blocks using a block partition mode set that combines one or a plurality of block partition modes that define a partition type, and decodes the plurality of blocks, wherein the block partition mode set includes a first block partition mode that defines a partition direction and a partition number for partitioning a first block, and a second block partition mode that defines a partition direction and a partition number for partitioning a second block, which is one of the blocks obtained after partitioning the first block, and in the partitioning, if the partition number of the first block partition mode is 3, the second block is a central block among the blocks obtained after partitioning the first block, and the partition direction of the second block partition mode is the same as the partition direction of the first block partition mode, the second block partition mode may only include a block partition mode with the partition number of 3.

[0324] The parameters for identifying the second block division mode in the decoding method of an embodiment of the present disclosure may include a first flag indicating whether the block is to be divided horizontally or vertically, and may not include a second flag indicating the number of divisions into which the block is to be divided.

[0325] A decoding method of an embodiment of the present disclosure includes the steps of: dividing an encoded signal read from a memory into a plurality of blocks using a block partition mode set that combines one or a plurality of block partition modes that define a partition type; and decoding the plurality of blocks, wherein the block partition mode set includes a first block partition mode that defines a partition direction and a partition number for dividing a first block, and a second block partition mode that defines a partition direction and a partition number for dividing a second block, which is one of the blocks obtained after dividing the first block, and wherein the partitioning step does not require the use of the second block partition mode with the partition number of 2 when the partition number of the first block partition mode is 3, the second block is a central block among the blocks obtained after dividing the first block, and the partition direction of the second block partition mode is the same as the partition direction of the first block partition mode.

[0326] A decoding method of an embodiment of the present disclosure includes a step of dividing an encoded signal read from a memory into a plurality of blocks using a block partitioning mode set that combines one or a plurality of block partitioning modes that define a partition type, and a step of decoding the plurality of blocks, wherein the block partitioning mode set includes a first block partitioning mode and a second block partitioning mode that each define a partitioning direction and a partitioning number, and the dividing step may restrict the use of the second block partitioning mode in which the partitioning number is 2.

[0327] A decoding method of an embodiment of the present disclosure includes a step of dividing an encoded signal read from a memory into a block set consisting of a plurality of blocks using a block partitioning mode set that combines one or a plurality of block partitioning modes that define a partitioning type, and a step of decoding the plurality of blocks, and in the dividing step, when a first block set obtained using a first block partitioning mode set and a second block set obtained using a second block partitioning mode set are identical, the dividing step may include dividing using only either the first block partitioning mode set or the second block partitioning mode set.

[0328] A picture compression program of an embodiment of the present disclosure divides a picture read from a memory into a plurality of blocks using a block partition mode set that combines one or a plurality of block partition modes that define a partition type, and decodes the plurality of blocks, wherein the block partition mode set includes a first block partition mode that defines a partition direction and a partition number for partitioning a first block, and a second block partition mode that defines a partition direction and a partition number for partitioning a second block, which is one of the blocks obtained after partitioning the first block, and in the partitioning, if the partition number of the first block partition mode is 3, the second block is a central block among the blocks obtained after partitioning the first block, and the partition direction of the second block partition mode is the same as the partition direction of the first block partition mode, the second block partition mode may only include a block partition mode with the partition number of 3.

[0329] A picture compression program of an embodiment of the present disclosure includes a step of dividing a picture read from a memory into a plurality of blocks using a block division mode set that combines one or a plurality of block division modes that define a division type, and a step of encoding the plurality of blocks, wherein the block division mode set includes a first block division mode that defines a division direction and a division number for dividing a first block, and a second block division mode that defines a division direction and a division number for dividing a second block, which is one of the blocks obtained after dividing the first block, and the division step does not need to use the second block division mode with the division number of 2 when the division number of the first block division mode is 3, the second block is a central block among the blocks obtained after dividing the first block, and the division direction of the second block division mode is the same as the division direction of the first block division mode.

[0330] A picture compression program of an embodiment of the present disclosure includes a step of dividing a picture read from a memory into a plurality of blocks using a block division mode set that combines one or a plurality of block division modes that define a division type, and a step of encoding the plurality of blocks, wherein the block division mode set includes a first block division mode and a second block division mode that each define a division direction and a division number, and the division step may restrict the use of the second block division mode in which the division number is 2.

[0331] A picture compression program of an embodiment of the present disclosure includes a step of dividing a picture read from a memory into a block set consisting of a plurality of blocks using a block partition mode set that combines one or a plurality of block partition modes that define a partition type, and a step of encoding the plurality of blocks, and in the dividing step, when a first block set obtained using a first block partition mode set and a second block set obtained using a second block partition mode set are identical, the dividing step may include dividing the picture using only either the first block partition mode set or the second block partition mode set. [Industrial Applicability]

[0332] The present invention can be used in encoding / decoding of multimedia data, particularly in image and video encoding / decoding devices that use block encoding / decoding. [Explanation of symbols]

[0333] 100 Encoding device 102 Division 104 Subtraction section 106, 5001 Converter 108, 5002 Quantization section 110, 5009 Entropy coding unit 112, 5003, 6002 Inverse quantization section 114, 5004, 6003 Reverse conversion unit 116 Addition section 118, 5005, 6004 block memory 120 Loop filter section 122, 5006, 6005 frame memory 124, 5007, 6006 Intra prediction section 126, 5008, 6007 Inter Prediction 128 Predictive control unit 200 Decryption device 202, 6001 Entropy Decoding Unit 204 Inverse quantization section 206 Inverse conversion unit 208 Addition section 210 Block Memory 212 Loop filter section 214 Frame Memory 216 Intra Prediction Unit 218 Inter Prediction Unit 220 Predictive control unit 5000 Video Encoding Device 5010, 6008 Block division decision unit 6000 Video Decoding Device

Claims

1. An encoding device for encoding a picture, comprising: A processor; A memory, The processor, in operation, Retrieving a block from a coding tree unit (CTU); a block division determination unit that divides the acquired block into a plurality of sub-blocks using one or a block division mode set that is a combination of a plurality of block division modes that define a division type; an encoding unit that encodes the plurality of sub-blocks; having the block division mode set includes a first block division mode that defines a division direction and a division number for dividing a first block, and a second block division mode that defines a division direction and a division number for dividing a second block, which is one of the sub-blocks obtained after dividing the first block, the block division determination unit determines, when the number of divisions in the first block division mode is three, the second block is a central block among sub-blocks obtained after division of the first block, and the division direction in the second block division mode is the same as the division direction in the first block division mode, that the parameters for identifying the second block division mode include a first flag indicating whether the sub-blocks are to be divided in a horizontal direction or a vertical direction, and do not include a second flag indicating the number of divisions into which the sub-blocks are to be divided; Encoding device.

2. A decoding device for decoding an encoded signal, comprising: A processor; A memory, The processor, in operation, Retrieving a block from a coding tree unit (CTU); a block division determination unit that divides the acquired block into a plurality of sub-blocks using one or a block division mode set that is a combination of a plurality of block division modes that define a division type; a decoding unit configured to decode the plurality of sub-blocks; having the block division mode set includes a first block division mode that defines a division direction and a division number for dividing a first block, and a second block division mode that defines a division direction and a division number for dividing a second block, which is one of the sub-blocks obtained after dividing the first block, the block division determination unit determines, when the number of divisions in the first block division mode is three, the second block is a central block among sub-blocks obtained after division of the first block, and the division direction in the second block division mode is the same as the division direction in the first block division mode, that the parameters for identifying the second block division mode include a first flag indicating whether the sub-blocks are to be divided in a horizontal direction or a vertical direction, and do not include a second flag indicating the number of divisions into which the sub-blocks are to be divided; Decryption device.

3. A transmitting device for transmitting a bit stream, comprising: A processor; A memory, The processor, in operation, Retrieving a block from a coding tree unit (CTU); a block division determination unit that divides the acquired block into a plurality of sub-blocks using one or a block division mode set that is a combination of a plurality of block division modes that define a division type; an encoding unit for encoding the plurality of sub-blocks into a bitstream; a transmitter for transmitting the encoded bitstream; having the block division mode set includes a first block division mode that defines a division direction and a division number for dividing a first block, and a second block division mode that defines a division direction and a division number for dividing a second block, which is one of the sub-blocks obtained after dividing the first block, the block division determination unit determines, when the number of divisions in the first block division mode is three, the second block is a central block among sub-blocks obtained after division of the first block, and the division direction in the second block division mode is the same as the division direction in the first block division mode, that the parameters for identifying the second block division mode include a first flag indicating whether the sub-blocks are to be divided in a horizontal direction or a vertical direction, and do not include a second flag indicating the number of divisions into which the sub-blocks are to be divided; Transmitting device.

4. A computer-readable non-transitory storage medium storing a transmission program, The transmission program, in its operation, Retrieving a block from a coding tree unit (CTU); Dividing the acquired block into a plurality of sub-blocks using one or a block division mode set including a combination of a plurality of block division modes that define a division type, and transmitting an encoded bitstream; the block division mode set includes a first block division mode that defines a division direction and a division number for dividing a first block, and a second block division mode that defines a division direction and a division number for dividing a second block, which is one of the sub-blocks obtained after dividing the first block, when the number of divisions in the first block division mode is three, the second block is a central block among sub-blocks obtained after division of the first block, and the division direction in the second block division mode is the same as the division direction in the first block division mode, the parameters for identifying the second block division mode include a first flag indicating whether the sub-blocks are divided in the horizontal direction or the vertical direction, and do not include a second flag indicating the number of divisions into which the sub-blocks are divided. Non-transitory storage medium.

Citation Information

Patent Citations

  • Methods and Apparatuses of Constrained Multi-type-tree Block Partition for Video Coding

    US20180103268A1

  • Method for transmitting and receiving channel state information in wireless communication system, and apparatus therefor

    WO2017090987A1

  • Multi-type-tree framework for video coding

    WO2017123980A1

  • Binary, ternary and QUAD tree partitioning for JVET coding of video data

    WO2017205700A1

  • Image decoding device and image coding device

    WO2018110600A1