Encoding device and decoding device
By combining the block segmentation mode set and optimizing the signalization of the block segmentation mode, the problem of image compression efficiency decline caused by the increase in signalization overhead of block segmentation information is solved, and more efficient block segmentation encoding is achieved.
Patent Information
- Application Number
- CN202210417607.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-02-20
- Filing Date
- 2019-05-09
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2039-05-09
AI Technical Summary
In the prior art, the signaling overhead of block segmentation information increases with the increase in the segmentation depth, resulting in a decrease in image compression efficiency.
Using an encoding device and a decoding device, block segmentation is performed by combining the block segmentation mode set, the first block segmentation mode and the second block segmentation mode are used to define the division direction and number, and unnecessary signalization is reduced by identifying the parameters of the second block segmentation mode. For example, when the division number of the first block segmentation mode is 3 and the division direction is the same, only the block segmentation mode with the division number of 3 is included.
It improves the encoding efficiency of block segmentation information, reduces signaling overhead, and improves image compression efficiency.
Smart Images

Figure CN114630115B_ABST
Abstract
Description
[0001] This application is a divisional of the invention patent application with the application date of May 9, 2019, the application number of 201980033751.4, and the invention title of "Coding Device, Decoding Device, Coding Method, Decoding Method, and Image Compression Program". Technical Field
[0002] The present disclosure relates to methods and apparatuses for encoding and decoding video and images using block partitioning. Background Art
[0003] In conventional image and video coding methods, an image is generally divided into blocks and encoded and decoded at the block level. In recent video standard development, in addition to typical sizes of 8×8 or 16×16, encoding and decoding processes can be performed with various block sizes. For image encoding and decoding processes, a range of sizes from 4×4 to 256×256 can be used.
[0004] Prior Art Documents
[0005] Non-Patent Documents
[0006] Non-Patent Document 1: H.265 (ISO / IEC 23008-2 HEVC (High Efficiency Video Coding)) Summary of the Invention
[0007] Problems to be Solved by the Invention
[0008] In order to represent a range of sizes from 4×4 to 256×256, block partitioning information such as block partitioning patterns (e.g., quad-tree, binary tree, and ternary tree) and partitioning flags (e.g., split flag) is determined and signaled for use in blocks. The overhead of this signaling increases as the depth of the partitioning increases. Also, the increased overhead reduces the video compression efficiency.
[0009] Therefore, an encoding device according to one aspect of the present disclosure provides an encoding device and the like capable of improving the compression efficiency in the encoding of block partitioning information.
[0010] Means for Solving the Problems
[0011] An encoding device according to an aspect of the present disclosure encodes an image, and includes: a processor; and a memory; the processor has: a block division determination unit that divides the image read from the memory into a plurality of blocks using a block division mode set obtained by combining one or more block division modes, the block division mode defining a division type; and an encoding unit that encodes the plurality of blocks; the block division mode set includes a first block division mode and a second block division mode, the first block division mode defining a division direction and a division number for dividing a first block, the second block division mode defining a division direction and a division number for dividing a second block that is one of the blocks obtained after dividing the first block; a parameter for identifying the second block division mode includes a first flag indicating in which direction, among the horizontal direction and the vertical direction, the block is divided; in the block division determination unit, when the division number of the first block division mode is 3, the second block is the central block among the blocks obtained after dividing the first block, and the division direction of the second block division mode indicated by the first flag is the same as the division direction of the first block division mode, the second block division mode includes only a block division mode with a division number of 3; when the division number of the first block division mode is 3, the second block is the central block among the blocks obtained after dividing the first block, and the division direction of the second block division mode indicated by the first flag is different from the division direction of the first block division mode, the second block division mode includes a block division mode with a division number of 2.
[0012] A decoding apparatus according to an aspect of the present disclosure decodes an encoded signal, and includes: a processor; and a memory. The processor includes: a block division determination unit that divides the encoded signal read from the memory into a plurality of blocks using a set of block division patterns obtained by combining one or more block division patterns, where the block division patterns define division types; and a decoding unit that decodes the plurality of blocks. The set of block division patterns includes a first block division pattern and a second block division pattern. The first block division pattern defines a division direction and a division number for dividing a first block, and the second block division pattern defines a division direction and a division number for dividing a second block, which is one of the blocks obtained by dividing the first block. A parameter for identifying the second block division pattern includes a first flag indicating in which direction, the horizontal direction or the vertical direction, the block is divided. In the block division determination unit, when the division number of the first block division pattern is 3, the second block is the central block among the blocks obtained by dividing the first block, and the division direction of the second block division pattern indicated by the first flag is the same as the division direction of the first block division pattern, the second block division pattern includes only a block division pattern with a division number of 3. When the division number of the first block division pattern is 3, the second block is the central block among the blocks obtained by dividing the first block, and the division direction of the second block division pattern indicated by the first flag is different from the division direction of the first block division pattern, the second block division pattern includes a block division pattern with a division number of 2.
[0013] In addition, these inclusive or specific forms can also be implemented by a system, a method, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or can be implemented by any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.
[0014] Advantageous Effects of the Invention
[0015] According to the present invention, it is possible to improve the compression efficiency in the encoding of block division information. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a block diagram showing the functional configuration of an encoding apparatus according to Embodiment 1.
[0017] Figure 2 is a diagram showing an example of block division in Embodiment 1.
[0018] Figure 3 is a table showing transform basis functions corresponding to respective transform types.
[0019] Figure 4AIt is a diagram showing an example of the shape of the filter used in the ALF.
[0020] Figure 4B It is a diagram showing another example of the shape of the filter used in the ALF.
[0021] Figure 4C It is a diagram showing another example of the shape of the filter used in the ALF.
[0022] Figure 5A It is a diagram showing 67 intra prediction modes of intra prediction.
[0023] Figure 5B It is a flowchart for explaining the outline of the predicted image correction process based on the OBMC process.
[0024] Figure 5C It is a conceptual diagram for explaining the outline of the predicted image correction process based on the OBMC process.
[0025] Figure 5D It is a diagram showing an example of FRUC.
[0026] Figure 6 It is a diagram for explaining pattern matching (bidirectional matching) between two blocks along a motion trajectory.
[0027] Figure 7 It is a diagram for explaining pattern matching (template matching) between a template within the current picture and a block within the reference picture.
[0028] Figure 8 It is a diagram for explaining a model assuming uniform linear motion.
[0029] Figure 9A It is a diagram for explaining the derivation of a motion vector in units of sub-blocks based on motion vectors of multiple adjacent blocks.
[0030] Figure 9B It is a diagram for explaining the outline of the motion vector derivation process based on the merge mode.
[0031] Figure 9C It is a conceptual diagram for explaining the outline of the DMVR process.
[0032] Figure 9D It is a diagram for explaining the outline of the predicted image generation method that employs a luminance correction process based on the LIC process.
[0033] Figure 10 It is a block diagram showing the functional structure of the decoding device related to Embodiment 1.
[0034] Figure 11 It is a flowchart showing the video encoding process related to Embodiment 2.
[0035] Figure 12 It is a flowchart showing the video decoding process related to Embodiment 2.
[0036] Figure 13 It is a flowchart showing the video encoding process related to Embodiment 3.
[0037] Figure 14 It is a flowchart showing the video decoding process related to Embodiment 3.
[0038] Figure 15 It is a block diagram showing the structure of a video / image encoding device related to Embodiment 2 or 3.
[0039] Figure 16 It is a block diagram showing the structure of a video / image decoding device related to Embodiment 2 or 3.
[0040] Figure 17 It is a diagram showing an example of the conceivable positions of the first parameter in the compressed video stream in Embodiment 2 or 3.
[0041] Figure 18 It is a diagram showing an example of the conceivable positions of the second parameter in the compressed video stream in Embodiment 2 or 3.
[0042] Figure 19 It is a diagram showing an example of the second parameter following the first parameter in Embodiment 2 or 3.
[0043] Figure 20 It is a diagram showing an example of not selecting the second block mode for the division of a 2N×N pixel block as shown in step (2c) in Embodiment 2.
[0044] Figure 21 It is a diagram showing an example of not selecting the second block mode for the division of an N×2N pixel block as shown in step (2c) in Embodiment 2.
[0045] Figure 22 It is a diagram showing an example of not selecting the second block mode for the division of an N×N pixel block as shown in step (2c) in Embodiment 2.
[0046] Figure 23 It is a diagram showing an example of not selecting the second block mode for the division of an N×N pixel block as shown in step (2C) in Embodiment 2.
[0047] Figure 24 It is a diagram showing an example of dividing a 2N×N pixel block using the block mode selected when not selecting the second block mode as shown in step (3) in Embodiment 2.
[0048] Figure 25 This is a diagram showing an example of dividing a block of N×2N pixels using the block division mode selected when the second block division mode is not selected as shown in step (3) in Embodiment 2.
[0049] Figure 26 This is a diagram showing an example of dividing a block of N×N pixels using the block division mode selected when the second block division mode is not selected as shown in step (3) in Embodiment 2.
[0050] Figure 27 This is a diagram showing an example of dividing a block of N×N pixels using the block division mode selected when the second block division mode is not selected as shown in step (3) in Embodiment 2.
[0051] Figure 28 This is a diagram showing an example of the block division mode for dividing a block of N×N pixels in Embodiment 2. Figure 28 (a) to (h) thereof are diagrams showing mutually different block division modes.
[0052] Figure 29 This is a diagram showing an example of the block division type and block division direction for dividing a block of N×N pixels in Embodiment 3. (1), (2), (3), and (4) are different block division types, (1a), (2a), (3a), and (4a) are block division modes with different block division types in the vertical block division direction, and (1b), (2b), (3b), and (4b) are block division modes with different block division types in the horizontal block division direction.
[0053] Figure 30 This is a diagram showing the advantages brought about by encoding the block division type before the block division direction as compared to encoding the block division direction before the block division type in Embodiment 3.
[0054] Figure 31A This is a diagram showing an example of dividing a block into sub-blocks using a block division mode set that uses fewer binary numbers in the encoding of the block division mode.
[0055] Figure 31B This is a diagram showing an example of dividing a block into sub-blocks using a block division mode set that uses fewer binary numbers in the encoding of the block division mode.
[0056] Figure 32A This is a diagram showing an example of dividing a block into sub-blocks using the block division mode set that first appears in a specified order of multiple block division mode sets.
[0057] Figure 32B This is a diagram showing an example of dividing a block into sub-blocks using the block division mode set that first appears in a specified order of multiple block division mode sets.
[0058] Figure 32C This is a diagram showing an example of dividing a block into sub - blocks by indicating the block pattern set that first appears in a specified order among multiple block pattern sets.
[0059] Figure 33 This is an overall structural diagram of a content supply system that implements a content distribution service.
[0060] Figure 34 This is a diagram showing an example of the coding structure in scalable coding.
[0061] Figure 35 This is a diagram showing an example of the coding structure in scalable coding.
[0062] Figure 36 This is a diagram showing an example of the display screen of a web page.
[0063] Figure 37 This is a diagram showing an example of the display screen of a web page.
[0064] Figure 38 This is a diagram showing an example of a smart phone.
[0065] Figure 39 This is a block diagram showing an example of the structure of a smart phone.
[0066] Figure 40 This is a diagram showing an example of the constraints of a block - splitting pattern that divides a rectangular block into 3 sub - blocks.
[0067] Figure 41 This is a diagram showing an example of the constraints of a block - splitting pattern that divides a block into 2 sub - blocks.
[0068] Figure 42 This is a diagram showing an example of the constraints of a block - splitting pattern that divides a square block into 3 sub - blocks.
[0069] Figure 43 This is a diagram showing an example of the constraints of a block - splitting pattern that divides a rectangular block into 2 sub - blocks.
[0070] Figure 44 This is a diagram showing an example of the constraints based on the splitting direction of a block - splitting pattern that divides a non - rectangular block into 2 sub - blocks.
[0071] Figure 45 This is a diagram showing an example of an effective splitting direction for splitting a non - rectangular block into 2 sub - blocks. Detailed implementation mode
[0072] Hereinafter, the implementation mode will be specifically described with reference to the accompanying drawings.
[0073] In addition, the embodiments described below all represent inclusive or specific examples. The numerical values, shapes, materials, components, configurations and connection forms of the components, steps, order of steps, etc. shown in the following embodiments are examples and do not limit the meaning of the claims. In addition, among the components of the following embodiments, the components not described in the independent claims representing the most general concepts are described as optional components.
[0074] (Embodiment 1)
[0075] First, as an example of an encoding device and a decoding device that can apply the processing and / or structure described in each form of the present invention described later, the outline of Embodiment 1 will be described. However, Embodiment 1 is merely an example of an encoding device and a decoding device that can apply the processing and / or structure described in each form of the present invention, and the processing and / or structure described in each form of the present invention can also be implemented in encoding devices and decoding devices different from Embodiment 1.
[0076] When applying the processing and / or structure described in each form of the present invention to Embodiment 1, for example, one of the following may be performed.
[0077] (1) For the encoding device or decoding device of Embodiment 1, replace the components corresponding to the components described in each form of the present invention among the multiple components constituting the encoding device or decoding device with the components described in each form of the present invention;
[0078] (2) For the encoding device or decoding device of Embodiment 1, after arbitrarily changing the addition, replacement, deletion, etc. of the functions or processes implemented on a part of the multiple components constituting the encoding device or decoding device, replace the components corresponding to the components described in each form of the present invention with the components described in each form of the present invention;
[0079] (3) For the method implemented by the encoding device or decoding device of Embodiment 1, add processing, and / or arbitrarily change the replacement, deletion, etc. of a part of the multiple processes included in the method, and then replace the process corresponding to the process described in each form of the present invention with the process described in each form of the present invention;
[0080] (4) Combine a part of the multiple components constituting the encoding device or decoding device of Embodiment 1 with the components described in each form of the present invention, the components having a part of the functions of the components described in each form of the present invention, or the components implementing a part of the processes implemented by the components described in each form of the present invention and implement;
[0081] (5) Combine an element that has part of the functions of part of the multiple elements constituting the encoding device or decoding device of Embodiment 1, or an element that implements part of the processing implemented by part of the multiple elements constituting the encoding device or decoding device of Embodiment 1, with the elements described in each aspect of the present invention, an element that has part of the functions of the elements described in each aspect of the present invention, or an element that implements part of the processing implemented by the elements described in each aspect of the present invention;
[0082] (6) For the method implemented by the encoding device or decoding device of Embodiment 1, replace the processing corresponding to the processing described in each aspect of the present invention among the multiple processes included in the method with the processing described in each aspect of the present invention;
[0083] (7) Combine part of the processes included in the method implemented by the encoding device or decoding device of Embodiment 1 with the processing described in each aspect of the present invention.
[0084] In addition, the implementation manners of the processing and / or structure described in each aspect of the present invention are not limited to the above examples. For example, it can also be implemented in a device used for a purpose different from the moving image / image encoding device or moving image / image decoding device disclosed in Embodiment 1, and the processing and / or structure described in each aspect can also be implemented independently. In addition, the processing and / or structure described in different aspects can also be combined and implemented.
[0085] [Outline of Encoding Device]
[0086] First, the outline of the encoding device of Embodiment 1 will be described. Figure 1 It is a block diagram showing the functional structure of the encoding device 100 of Embodiment 1. The encoding device 100 is a moving image / image encoding device that encodes moving images / images in units of blocks.
[0087] As Figure 1 shown, the encoding device 100 is a device that encodes images in units of blocks, and includes a division unit 102, a subtraction unit 104, a transformation unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse transformation unit 114, an addition unit 116, a block memory 118, a loop filter unit 120, a frame memory 122, an intra prediction unit 124, an inter prediction unit 126, and a prediction control unit 128.
[0088] The encoding device 100 is implemented by, for example, a general-purpose processor and a memory. In this case, when the software program stored in the memory is executed by the processor, the processor functions as a splitting unit 102, a subtraction unit 104, a transformation unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse transformation unit 114, an addition unit 116, a loop filtering unit 120, an intra prediction unit 124, an inter prediction unit 126, and a prediction control unit 128. In addition, the encoding device 100 can also be implemented as one or more dedicated electronic circuits corresponding to the splitting unit 102, the subtraction unit 104, the transformation unit 106, the quantization unit 108, the entropy encoding unit 110, the inverse quantization unit 112, the inverse transformation unit 114, the addition unit 116, the loop filtering unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128.
[0089] Hereinafter, each component included in the encoding device 100 will be described.
[0090] [Splitting Unit]
[0091] The splitting unit 102 splits each picture included in the input moving image into a plurality of blocks and outputs each block to the subtraction unit 104. For example, the splitting unit 102 first splits the picture into blocks of a fixed size (e.g., 128×128). Such blocks of a fixed size are sometimes referred to as coding tree units (CTUs). And the splitting unit 102 splits each block of a fixed size into blocks of a variable size (e.g., 64×64 or less) based on recursive quadtree and / or binary tree block splitting. Such blocks of a variable size are sometimes referred to as coding units (CUs), prediction units (PUs), or transformation units (TUs). Additionally, in the present embodiment, it is not necessary to distinguish between CUs, PUs, and TUs, and a part or all of the blocks in the picture can also be used as the processing units for CUs, PUs, and TUs.
[0092] Figure 2 is a diagram showing an example of block splitting in Embodiment 1. In Figure 2 the solid lines represent block boundaries based on quadtree block splitting, and the dashed lines represent block boundaries based on binary tree block splitting.
[0093] Here, block 10 is a square block of 128×128 pixels (128×128 block). This 128×128 block 10 is first split into four square 64×64 blocks (quadtree block splitting).
[0094] The 64×64 block in the upper left is further vertically divided into two rectangular 32×64 blocks, and the 32×64 block on the left is further vertically divided into two rectangular 16×64 blocks (binary tree block division). As a result, the 64×64 block in the upper left is divided into two 16×64 blocks 11, 12 and a 32×64 block 13.
[0095] The 64×64 block in the upper right is horizontally divided into two rectangular 64×32 blocks 14, 15 (binary tree block division).
[0096] The 64×64 block in the lower left is divided into four square 32×32 blocks (quadtree block division). The upper left block and the lower right block among the four 32×32 blocks are further divided. The upper left 32×32 block is vertically divided into two rectangular 16×32 blocks, and the right 16×32 block is further horizontally divided into two 16×16 blocks (binary tree block division). The lower right 32×32 block is horizontally divided into two 32×16 blocks (binary tree block division). As a result, the 64×64 block in the lower left is divided into a 16×32 block 16, two 16×16 blocks 17, 18, two 32×32 blocks 19, 20, and two 32×16 blocks 21, 22.
[0097] The 64×64 block 23 in the lower right is not divided.
[0098] As described above, in Figure 2 , the block 10 is divided into 13 variable-sized blocks 11 to 23 based on recursive quadtree and binary tree block division. Such a division is sometimes called QTBT (quad - tree plus binary tree) division.
[0099] In addition, in Figure 2 , one block is divided into four or two blocks (quadtree or binary tree block division), but the division is not limited to this. For example, one block can also be divided into three blocks (ternary tree division). The division including such a ternary tree division is sometimes called MBT (multi type tree) division.
[0100] [Subtraction section]
[0101] The subtraction section 104 subtracts the predicted signal (predicted sample) from the original signal (original sample) in units of the blocks divided by the division section 102. That is, the subtraction section 104 calculates the prediction error (also called the residual) of the block to be encoded (hereinafter referred to as the current block). And the subtraction section 104 outputs the calculated prediction error to the transformation section 106.
[0102] The original signal is the input signal of the encoding device 100, and is a signal representing the images of each picture constituting the moving image (for example, a luma signal and two chroma signals). Hereinafter, the signal representing the image may also be referred to as a sample.
[0103] [Transform section]
[0104] The transform section 106 transforms the prediction error in the spatial domain into transform coefficients in the frequency domain, and outputs the transform coefficients to the vectorization section 108. Specifically, the transform section 106 performs a preset discrete cosine transform (DCT) or discrete sine transform (DST) on the prediction error in the spatial domain, for example.
[0105] In addition, the transform section 106 may adaptively select a transform type from among multiple transform types, and use a transform basis function corresponding to the selected transform type to transform the prediction error into transform coefficients. Such a transform is sometimes referred to as EMT (explicit multiple core transform) or AMT (adaptive multiple transform).
[0106] The multiple transform types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. Figure 3 is a table showing the transform basis functions corresponding to the respective transform types. In Figure 3 N represents the number of input pixels. The selection of the transform type from among these multiple transform types may depend, for example, on the type of prediction (intra prediction and inter prediction), or on the intra prediction mode.
[0107] Information indicating whether to apply such EMT or AMT (for example, referred to as an AMT flag) and information indicating the selected transform type are signaled at the CU level. In addition, the signaling of this information does not need to be limited to the CU level, and may also be at other levels (for example, sequence level, picture level, slice level, tile level, or CTU level).
[0108] In addition, the transformation unit 106 can also perform inverse transformation on the transformation coefficients (transformation results). Such inverse transformation is sometimes referred to as AST (adaptive secondary transform) or NSST (non-separable secondary transform). For example, the transformation unit 106 performs inverse transformation on each sub-block (e.g., 4×4 sub-block) included in the block of transformation coefficients corresponding to the intra-prediction error. Information indicating whether to apply NSST and information related to the transformation matrix used in NSST are signaled at the CU level. In addition, the signaling of this information does not need to be limited to the CU level and can also be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).
[0109] Here, a Separable transformation refers to a method of performing multiple transformations by separating them in each direction according to the number of dimensions of the input, and a Non-Separable transformation refers to a method of treating two or more dimensions as one dimension and performing transformation together when the input is multi-dimensional.
[0110] For example, as an example of a Non-Separable transformation, when the input is a 4×4 block, it can be regarded as a permutation with 16 elements, and a transformation process is performed on this permutation with a 16×16 transformation matrix.
[0111] In addition, similarly, a method of performing Givens rotation on this permutation multiple times (Hypercube Givens Transform) after regarding a 4×4 input block as a permutation with 16 elements is also an example of a Non-Separable transformation.
[0112] [Quantization Unit]
[0113] The quantization unit 108 quantizes the transformation coefficients output from the transformation unit 106. Specifically, the quantization unit 108 scans the transformation coefficients of the current block in a prescribed scanning order, and quantizes the transformation coefficients based on the quantization parameter (QP) corresponding to the scanned transformation coefficients. And the quantization unit 108 outputs the quantized transformation coefficients (hereinafter referred to as quantization coefficients) of the current block to the entropy coding unit 110 and the inverse quantization unit 112.
[0114] The prescribed order is the order for quantization / inverse quantization of the transformation coefficients. For example, the prescribed scanning order is defined by ascending frequency (order from low frequency to high frequency) or descending frequency (order from high frequency to low frequency).
[0115] A quantization parameter refers to a parameter that defines the quantization step (quantization width). For example, if the value of the quantization parameter increases, the quantization step also increases. That is, if the value of the quantization parameter increases, the quantization error increases.
[0116] [Entropy Encoding Unit]
[0117] The entropy encoding unit 110 generates an encoded signal (encoded bitstream) by performing variable-length encoding on the quantized coefficients that are the input from the quantization unit 108. Specifically, the entropy encoding unit 110 binarizes the quantized coefficients, for example, and performs arithmetic coding on the binary signal.
[0118] [Inverse Quantization Unit]
[0119] The inverse quantization unit 112 performs inverse quantization on the quantized coefficients that are the input from the quantization unit 108. Specifically, the inverse quantization unit 112 performs inverse quantization on the quantized coefficients of the current block in a prescribed scan order. And the inverse quantization unit 112 outputs the inverse-quantized transform coefficients of the current block to the inverse transform unit 114.
[0120] [Inverse Transform Unit]
[0121] The inverse transform unit 114 restores the prediction error by performing an inverse transform on the transform coefficients that are the input from the inverse quantization unit 112. Specifically, the inverse transform unit 114 restores the prediction error of the current block by performing an inverse transform corresponding to the transform of the transform unit 106 on the transform coefficients. And the inverse transform unit 114 outputs the restored prediction error to the addition unit 116.
[0122] In addition, since information is lost due to quantization in the restored prediction error, it is not consistent with the prediction error calculated by the subtraction unit 104. That is, the restored prediction error contains quantization error.
[0123] [Addition Unit]
[0124] The addition unit 116 reconstructs the current block by adding the prediction error that is the input from the inverse transform unit 114 and the prediction sample that is the input from the prediction control unit 128. And the addition unit 116 outputs the reconstructed block to the block memory 118 and the loop filter unit 120. The reconstructed block is sometimes called a local decoded block.
[0125] [Block Memory]
[0126] The block memory 118 is a storage unit for storing blocks within the picture to be coded (hereinafter referred to as the current picture) that are referred to in intra prediction. Specifically, the block memory 118 stores the reconstructed block output from the addition unit 116.
[0127] [Loop Filter Unit]
[0128] The loop filter unit 120 performs loop filtering on the block reconstructed by the adder unit 116, and outputs the filtered reconstructed block to the frame memory 122. Loop filtering refers to filtering used within the coding loop (in-loop filtering), and includes, for example, deblocking filtering (DF), sample adaptive offset (SAO), and adaptive loop filtering (ALF).
[0129] In ALF, a least squares error filter for removing coding distortion is employed. For example, for each 2×2 sub-block within the current block, one filter selected from multiple filters is used based on the direction and activity of the locality-based gradient.
[0130] Specifically, first, sub-blocks (e.g., 2×2 sub-blocks) are classified into multiple classes (e.g., 15 or 25 classes). The classification of sub-blocks is performed based on the direction and activity of the gradient. For example, using the direction value D of the gradient (e.g., 0 to 2 or 0 to 4) and the activity value A of the gradient (e.g., 0 to 4), the classification value C is calculated (e.g., C = 5D + A). And based on the classification value C, the sub-blocks are classified into multiple classes (e.g., 15 or 25 classes).
[0131] The direction value D of the gradient is derived, for example, by comparing the gradients in multiple directions (e.g., horizontal, vertical, and two diagonal directions). In addition, the activity value A of the gradient is derived, for example, by adding the gradients in multiple directions and quantifying the added result.
[0132] Based on the result of such classification, the filter for the sub-block is determined from among multiple filters.
[0133] As the shape of the filter used in ALF, for example, a circularly symmetric shape is used. Figures 4A to 4C It is a diagram showing multiple examples of the shape of the filter used in ALF. Figure 4A Shows a 5×5 diamond-shaped filter, Figure 4B Shows a 7×7 diamond-shaped filter, Figure 4C Shows a 9×9 diamond-shaped filter. The information indicating the shape of the filter is signaled at the picture level. Additionally, the signaling of the information indicating the shape of the filter does not need to be limited to the picture level and can also be at other levels (e.g., sequence level, slice level, tile level, CTU level, or CU level).
[0134] The on / off of ALF is determined, for example, at the picture level or CU level. For example, for luminance, it is determined at the CU level whether to employ ALF, and for chrominance, it is determined at the picture level whether to employ ALF. The information indicating the on / off of ALF is signaled at the picture level or CU level. Additionally, the signaling of the information indicating the on / off of ALF does not need to be limited to the picture level or CU level and can also be at other levels (e.g., sequence level, slice level, tile level, or CTU level).
[0135] Coefficient sets of multiple selectable filters (e.g., up to 15 or 25 filters) are signaled at the picture level. In addition, the signaling of the coefficient sets does not need to be limited to the picture level and can also be other levels (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).
[0136] [Frame memory]
[0137] The frame memory 122 is a storage unit for storing reference pictures used in inter-frame prediction, and is sometimes referred to as a frame buffer. Specifically, the frame memory 122 stores the reconstructed blocks filtered by the loop filter unit 120.
[0138] [Intra-frame prediction unit]
[0139] The intra-frame prediction unit 124 performs intra-frame prediction (also referred to as intra-picture prediction) of the current block by referring to the block in the current picture stored in the block memory 118, thereby generating a prediction signal (intra-frame prediction signal). Specifically, the intra-frame prediction unit 124 generates an intra-frame prediction signal by performing intra-frame prediction with reference to samples (e.g., luminance values, chrominance difference values) of blocks adjacent to the current block, and outputs the intra-frame prediction signal to the prediction control unit 128.
[0140] For example, the intra-frame prediction unit 124 performs intra-frame prediction using one of a plurality of predefined intra-frame prediction modes. The plurality of intra-frame prediction modes include one or more non-directional prediction modes and a plurality of directional prediction modes.
[0141] One or more non-directional prediction modes include, for example, the Planar (plane) prediction mode and the DC prediction mode defined by the H.265 / HEVC (High-Efficiency Video Coding) standard (Non-Patent Document 1).
[0142] The plurality of directional prediction modes include, for example, 33-direction prediction modes defined by the H.265 / HEVC standard. In addition, the plurality of directional prediction modes may include 32-direction prediction modes (a total of 65 directional prediction modes) in addition to the 33 directions. Figure 5A It is a diagram showing 67 intra-frame prediction modes (2 non-directional prediction modes and 65 directional prediction modes) in intra-frame prediction. The solid arrows indicate 33 directions defined by the H.265 / HEVC standard, and the dashed arrows indicate the additional 32 directions.
[0143] In addition, in the intra prediction of the chrominance blocks, the luminance blocks may also be referred to. That is, the chrominance components of the current block may also be predicted based on the luminance component of the current block. Such intra prediction is called CCLM (cross-component linear model) prediction in some cases. The intra prediction mode of the chrominance blocks that refers to the luminance blocks (e.g., called the CCLM mode) may also be added as one of the intra prediction modes of the chrominance blocks.
[0144] The intra prediction unit 124 may also correct the pixel value after intra prediction based on the gradients of the reference pixels in the horizontal / vertical directions. The intra prediction accompanied by such correction is called PDPC (position dependent intraprediction combination) in some cases. The information indicating whether PDPC is used (e.g., called the PDPC flag) is signaled at the CU level, for example. In addition, the signaling of this information does not need to be limited to the CU level and may also be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).
[0145] [Inter prediction unit]
[0146] The inter prediction unit 126 performs inter prediction (also called inter-picture prediction) of the current block by referring to a reference picture different from the current picture stored in the frame memory 122, thereby generating a prediction signal (inter prediction signal). The inter prediction is performed in units of the current block or sub-blocks (e.g., 4×4 blocks) within the current block. For example, the inter prediction unit 126 performs motion estimation within the reference picture for the current block or sub-block. And the inter prediction unit 126 uses the motion information (e.g., motion vector) obtained by the motion estimation to perform motion compensation, thereby generating the inter prediction signal of the current block or sub-block. And the inter prediction unit 126 outputs the generated inter prediction signal to the prediction control unit 128.
[0147] The motion information used in the motion compensation is signaled. In the signaling of the motion vector, a motion vector predictor may also be used. That is, the difference between the motion vector and the predicted motion vector may also be signaled.
[0148] Alternatively, it is also possible to use not only the motion information of the current block obtained by motion estimation but also the motion information of adjacent blocks to generate an inter-frame prediction signal. Specifically, it is also possible to perform weighted addition of a prediction signal based on the motion information obtained by motion estimation and a prediction signal based on the motion information of adjacent blocks, thereby generating an inter-frame prediction signal in units of sub-blocks within the current block. Such inter-frame prediction (motion compensation) is sometimes referred to as OBMC (overlapped block motion compensation).
[0149] In such an OBMC mode, information indicating the size of the sub-blocks used for OBMC (e.g., referred to as the OBMC block size) is signaled at the sequence level. In addition, information indicating whether the OBMC mode is adopted (e.g., referred to as the OBMC flag) is signaled at the CU level. Additionally, the levels at which these pieces of information are signaled do not need to be limited to the sequence level and the CU level, and can also be other levels (e.g., picture level, slice level, tile level, CTU level, or sub-block level).
[0150] A more specific description of the OBMC mode will be given. Figure 5B and Figure 5C are a flowchart and a conceptual diagram for explaining the outline of the predicted image correction process based on OBMC processing.
[0151] First, using the motion vector (MV) assigned to the coding target block, a predicted image (Pred) obtained by normal motion compensation is acquired.
[0152] Next, a predicted image (Pred_L) is acquired by adopting the motion vector (MV_L) of the already-coded left adjacent block for the coding target block, and the first correction of the predicted image is performed by weighted superposition of the above-mentioned predicted image and Pred_L.
[0153] Similarly, a predicted image (Pred_U) is acquired by adopting the motion vector (MV_U) of the already-coded upper adjacent block for the coding target block, and the second correction of the predicted image is performed by weighted superposition of the predicted image after the first correction and Pred_U, and this is used as the final predicted image.
[0154] In addition, a method of two-stage correction using the left adjacent block and the upper adjacent block has been described here, but it can also be configured to perform more than two-stage corrections using the right adjacent block and the lower adjacent block.
[0155] In addition, the region for superposition may not be the entire pixel region of the block, but only a partial region near the block boundary.
[0156] In addition, the prediction image correction process based on one reference picture has been described here, but the same applies to the case of correcting the prediction image based on multiple reference pictures. After obtaining the corrected prediction images according to the respective reference pictures, the obtained prediction images are further superimposed to obtain the final prediction image.
[0157] In addition, the processing target block described above may be in units of prediction blocks or in units of sub-blocks obtained by further dividing the prediction blocks.
[0158] As a method for determining whether to use the OBMC process, for example, there is a method of using a signal indicating whether to use the OBMC process, namely, obmc_flag. As a specific example, in an encoding device, it is determined whether the encoding target block belongs to a region with complex motion. In the case where it belongs to a region with complex motion, the value 1 is set as obmc_flag and encoding is performed using the OBMC process. In the case where it does not belong to a region with complex motion, the value 0 is set as obmc_flag and encoding is performed without using the OBMC process. On the other hand, in a decoding device, decoding is performed by decoding the obmc_flag described in the stream and switching whether to use the OBMC process according to its value.
[0159] In addition, the motion information may not be signaled and may be derived on the decoding device side. For example, the merge mode defined by the H.265 / HEVC standard may be used. In addition, for example, the motion information may be derived by performing motion estimation on the decoding device side. In this case, motion estimation is performed without using the pixel values of the current block.
[0160] Here, the mode of performing motion estimation on the decoding device side is described. The mode of performing motion estimation on the decoding device side is sometimes referred to as the PMMVD (pattern matched motion vector derivation) mode or the FRUC (frame rate up-conversion) mode.
[0161] In Figure 5D shows an example of the FRUC process. First, with reference to the motion vectors of the encoded blocks adjacent to the current block in space or time, a plurality of candidates each having a predicted motion vector are generated (it may be shared with the merge list). Then, the best candidate MV is selected from among the plurality of candidate MVs registered in the candidate list. For example, the evaluation value of each candidate included in the candidate list is calculated, and one candidate is selected based on the evaluation value.
[0162] Further, based on the motion vector of the selected candidate, a motion vector for the current block is derived. Specifically, for example, the motion vector of the selected candidate (the best candidate MV) is directly derived as the motion vector for the current block. In addition, for example, a motion vector for the current block may also be derived by performing pattern matching in a peripheral region of the position within the reference picture corresponding to the motion vector of the selected candidate. That is, the peripheral region of the best candidate MV may be searched by the same method, and if there is an MV with a better evaluation value, the best candidate MV is updated to the above MV, and it is used as the final MV of the current block. Additionally, a structure that does not perform this process may also be implemented.
[0163] The exact same process may also be performed when processing is carried out in units of sub-blocks.
[0164] Furthermore, regarding the evaluation value, it is calculated by obtaining a difference value of the reconstructed image through pattern matching between the region within the reference picture corresponding to the motion vector and a specified region. Additionally, it may also be that, in addition to the difference value, other information is used to calculate the evaluation value.
[0165] As the pattern matching, the first pattern matching or the second pattern matching is used. In some cases, the first pattern matching and the second pattern matching are respectively referred to as bilateral matching and template matching.
[0166] In the first pattern matching, pattern matching is performed between two blocks along the motion trajectory of the current block in two different reference pictures. Therefore, in the first pattern matching, as the specified region for calculating the evaluation value of the candidate, a region within another reference picture along the motion trajectory of the current block is used.
[0167] Figure 6 It is a diagram for explaining an example of pattern matching (bilateral matching) between two blocks along a motion trajectory. As Figure 6 shown, in the first pattern matching, by searching for the most matching pair among pairs of two blocks in two different reference pictures (Ref0, Ref1) along the motion trajectory of the current block (Cur block), two motion vectors (MV0, MV1) are derived. Specifically, for the current block, the difference between the reconstructed image at the specified position within the first encoded reference picture (Ref0) specified by the candidate MV and the reconstructed image at the specified position within the second encoded reference picture (Ref1) specified by the symmetric MV obtained by scaling the above candidate MV by the display time interval is obtained, and the evaluation value is calculated using the obtained difference value. A candidate MV with the best evaluation value among multiple candidate MVs may be selected as the final MV.
[0168] Under the assumption of a continuous motion trajectory, the motion vectors (MV0, MV1) indicating two reference blocks are proportional to the temporal distances (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, when the current picture is temporally located between the two reference pictures and the temporal distances from the current picture to the two reference pictures are equal, in the first pattern matching, a mirror-symmetric bidirectional motion vector is derived.
[0169] In the second pattern matching, pattern matching is performed between a template within the current picture (a block adjacent to the current block within the current picture (e.g., an upper and / or left adjacent block)) and a block within the reference picture. Therefore, in the second pattern matching, as the specified region for calculating the evaluation value for the above candidate, a block adjacent to the current block within the current picture is used.
[0170] Figure 7 It is a diagram for explaining an example of pattern matching (template matching) between a template within the current picture and a block within the reference picture. As Figure 7 shown, in the second pattern matching, by searching within the reference picture (Ref0) for the block that best matches the block adjacent to the current block (Cur block) within the current picture (Cur Pic), the motion vector of the current block is derived. Specifically, for the current block, the difference between the reconstructed image of the encoded region of both or one of the left adjacent and upper adjacent blocks and the reconstructed image at the equivalent position within the encoded reference picture (Ref0) specified by the candidate MV is derived, and the obtained difference value is used to calculate the evaluation value. Among the multiple candidate MVs, the candidate MV with the best evaluation value is selected as the best candidate MV.
[0171] Information indicating whether to adopt the FRUC mode (e.g., called the FRUC flag) is signaled at the CU level. In addition, in the case of adopting the FRUC mode (e.g., when the FRUC flag is true), information indicating the method of pattern matching (the first pattern matching or the second pattern matching) (e.g., called the FRUC mode flag) is signaled at the CU level. Additionally, the signaling of this information does not need to be limited to the CU level and can also be other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).
[0172] Here, a mode of deriving a motion vector based on a model assuming uniform linear motion is described. This mode has a case called BIO (bi-directional optical flow).
[0173] Figure 8 It is a diagram for explaining a model assuming uniform linear motion. InFigure 8 In this case, (v x , v y ) represents the velocity vector, and τ0 and τ1 respectively represent the temporal distances between the current picture (Cur Pic) and two reference pictures (Ref0, Ref1). (MVx0, MVy0) represents the motion vector corresponding to the reference picture Ref0, and (MVx1, MVy1) represents the motion vector corresponding to the reference picture Ref1.
[0174] At this time, under the assumption of a uniform linear motion of the velocity vector (v x , v y ), (MVx0, MVy0) and (MVx1, MVy1) are respectively expressed as (v x τ0, v y τ0) and (-v x τ1, -v y τ1), and the following optical flow equation (1) holds.
[0175] [Equation 1]
[0176]
[0177] Here, I (k) represents the luminance value of the reference image k (k = 0, 1) after motion compensation. This optical flow equation means that the sum of (i) the temporal differential of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Based on the combination of this optical flow equation and Hermite interpolation, the motion vector in block units obtained from a merge list or the like is corrected in pixel units.
[0178] In addition, the motion vector can also be derived on the decoder side by a method different from the derivation of the motion vector based on a model assuming uniform linear motion. For example, the motion vector can also be derived in sub-block units based on the motion vectors of multiple adjacent blocks.
[0179] Here, a mode of deriving the motion vector in sub-block units based on the motion vectors of multiple adjacent blocks will be described. This mode is called the affine motion compensation prediction mode in some cases.
[0180] Figure 9A is a diagram for explaining the derivation of the motion vector in sub-block units based on the motion vectors of multiple adjacent blocks. In Figure 9AAmong them, the current block includes 16 4×4 sub-blocks. Here, based on the motion vectors of adjacent blocks, the motion vector v0 of the upper left control point of the current block is derived, and based on the motion vectors of adjacent sub-blocks, the motion vector v1 of the upper right control point of the current block is derived. And, using the two motion vectors v0 and v1, the motion vectors (v x , v y ) of each sub-block within the current block are derived through the following equation (2).
[0181] [Equation 2]
[0182]
[0183] Here, x and y respectively represent the horizontal position and vertical position of the sub-block, and w represents a preset weight coefficient.
[0184] In such an affine motion compensation prediction mode, several modes with different methods for deriving the motion vectors of the upper left and upper right control points may also be included. Information representing such an affine motion compensation prediction mode (for example, called an affine flag) is signaled at the CU level. In addition, the signaling of the information representing this affine motion compensation prediction mode does not need to be limited to the CU level, and may also be other levels (for example, sequence level, picture level, slice level, tile level, CTU level, or sub-block level).
[0185] [Prediction control unit]
[0186] The prediction control unit 128 selects one of the intra-prediction signal and the inter-prediction signal, uses the selected signal as the prediction signal, and outputs it to the subtraction unit 104 and the addition unit 116.
[0187] Here, an example of deriving the motion vector of the coded object picture through the merge mode is described. Figure 9B It is a diagram for explaining the outline of the motion vector derivation process based on the merge mode.
[0188] First, a prediction MV list registering candidates of the prediction MV is generated. As candidates of the prediction MV, there are the MV of multiple coded blocks spatially located around the coded object block, that is, the spatial neighboring prediction MV, the MV of the block near the projection of the position of the coded object block in the coded reference picture, that is, the temporal neighboring prediction MV, the MV generated by combining the MV values of the spatial neighboring prediction MV and the temporal neighboring prediction MV, that is, the combined prediction MV, and the MV with a value of zero, that is, the zero prediction MV, etc.
[0189] Next, by selecting one prediction MV from the multiple prediction MVs registered in the prediction MV list, it is determined as the MV of the coded object block.
[0190] Furthermore, in the variable-length coding section, the merge_idx, which is a signal indicating which predicted MV is selected, is described in the stream and encoded.
[0191] In addition, Figure 9B The predicted MVs registered in the predicted MV list described in is an example. The number may be different from the number in the figure, or it may be a structure that does not include some types of the predicted MVs in the figure, or a structure with predicted MVs other than the types of the predicted MVs in the figure added.
[0192] In addition, the MV of the coding target block derived through the merge mode may be used for the subsequent DMVR processing to determine the final MV.
[0193] Here, an example of determining the MV using the DMVR processing is described.
[0194] Figure 9C is a conceptual diagram for explaining the outline of the DMVR processing.
[0195] First, the optimal MVP set for the processing target block is used as the candidate MV. According to the above candidate MVs, reference pixels are obtained from the first reference picture of the processed pictures in the L0 direction and the second reference picture of the processed pictures in the L1 direction respectively, and a template is generated by taking the average of each reference pixel.
[0196] Next, using the above template, the peripheral areas of the candidate MVs in the first reference picture and the second reference picture are searched respectively, and the MV with the minimum cost is determined as the final MV. In addition, for the cost value, it is calculated using the difference values between the pixel values of the template and the pixel values of the search area and the MV value, etc.
[0197] In addition, in the encoding device and the decoding device, the outline of the processing described here is basically common.
[0198] In addition, even if it is not the processing itself described here, as long as it is a processing that can search the periphery of the candidate MV and derive the final MV, other processing can also be used.
[0199] Here, the mode of generating a predicted image using the LIC processing is described.
[0200] Figure 9D is a diagram for explaining the outline of the predicted image generation method using the luminance correction processing based on the LIC processing.
[0201] First, an MV for obtaining a reference image corresponding to the coding target block from the reference picture of the encoded pictures is derived.
[0202] Next, for the coded object block, using the luminance pixel values of the left and upper adjacent coded neighboring reference regions and the luminance pixel values at the same position in the reference picture specified by the MV, information indicating how the luminance values change in the reference picture and the coded object picture is extracted, and a luminance correction parameter is calculated.
[0203] By performing a luminance correction process on the reference image in the reference picture specified by the MV using the above luminance correction parameter, a prediction image for the coded object block is generated.
[0204] In addition, Figure 9D the shape of the above neighboring reference region in Figure 9D is an example, and shapes other than this can also be used.
[0205] Furthermore, the process of generating a prediction image based on one reference picture has been described here, but the same applies when generating a prediction image based on multiple reference pictures. After performing a luminance correction process on the reference images obtained from each reference picture in the same manner, a prediction image is generated.
[0206] As a method for determining whether to adopt the LIC process, for example, there is a method of using the lic_flag as a signal indicating whether to adopt the LIC process. As a specific example, in the encoding device, it is determined whether the coded object block belongs to a region where a luminance change has occurred. If it belongs to a region where a luminance change has occurred, the value 1 is set as the lic_flag, and encoding is performed using the LIC process. If it does not belong to a region where a luminance change has occurred, the value 0 is set as the lic_flag, and encoding is performed without using the LIC process. On the other hand, in the decoding device, by decoding the lic_flag described in the stream, decoding is performed by switching whether to adopt the LIC process according to its value.
[0207] As another method for determining whether to adopt the LIC process, for example, there is also a method of determining according to whether the LIC process has been adopted in neighboring blocks. As a specific example, in the case where the coded object block is in the merge mode, it is determined whether the neighboring coded blocks selected during the derivation of the MV in the merge mode process have been encoded using the LIC process, and encoding is performed by switching whether to adopt the LIC process according to the result. In addition, in this example, the process during decoding is exactly the same.
[0208] [Outline of Decoding Device]
[0209] Next, an outline of a decoding device that can decode the encoded signal (encoded bitstream) output from the above encoding device 100 will be described. Figure 10FIG. 0 is a block diagram showing the functional configuration of the decoding apparatus 200 according to Embodiment 1. The decoding apparatus 200 is a moving image / image decoding apparatus that decodes moving images / images in units of blocks.
[0210] As Figure 10 shown, the decoding apparatus 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transform unit 206, an addition unit 208, a block memory 210, a loop filter unit 212, a frame memory 214, an intra prediction unit 216, an inter prediction unit 218, and a prediction control unit 220.
[0211] The decoding apparatus 200 is implemented by, for example, a general-purpose processor and a memory. In this case, when a software program stored in the memory is executed by the processor, the processor functions as the entropy decoding unit 202, the inverse quantization unit 204, the inverse transform unit 206, the addition unit 208, the loop filter unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220. Further, the decoding apparatus 200 may be implemented as one or more dedicated electronic circuits corresponding to the entropy decoding unit 202, the inverse quantization unit 204, the inverse transform unit 206, the addition unit 208, the loop filter unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220.
[0212] Hereinafter, each component included in the decoding apparatus 200 will be described.
[0213] [Entropy Decoding Unit]
[0214] The entropy decoding unit 202 performs entropy decoding on the encoded bitstream. Specifically, the entropy decoding unit 202 arithmetically decodes the encoded bitstream into a binary signal, for example. Then, the entropy decoding unit 202 de-binarizes the binary signal. Thereby, the entropy decoding unit 202 outputs the quantization coefficients to the inverse quantization unit 204 in units of blocks.
[0215] [Inverse Quantization Unit]
[0216] The inverse quantization unit 204 performs inverse quantization on the quantization coefficients of the block to be decoded (hereinafter referred to as the current block) that is input from the entropy decoding unit 202. Specifically, the inverse quantization unit 204 performs inverse quantization on the quantization coefficients of the current block based on the quantization parameters corresponding to the quantization coefficients, respectively. And the inverse quantization unit 204 outputs the inverse-quantized quantization coefficients (i.e., transform coefficients) of the current block to the inverse transform unit 206.
[0217] [Inverse Transform Unit]
[0218] The inverse transform unit 206 restores the prediction error by performing inverse transform on the transform coefficients that are input from the inverse quantization unit 204.
[0219] For example, when the information decoded from the coded bitstream indicates the use of EMT or AMT (e.g., the AMT flag is true), the inverse transform unit 206 performs an inverse transform on the transform coefficients of the current block based on the decoded information indicating the transform type.
[0220] In addition, for example, when the information decoded from the coded bitstream indicates the use of NSST, the inverse transform unit 206 applies an inverse re - transform to the transform coefficients.
[0221] [Addition unit]
[0222] The addition unit 208 reconstructs the current block by adding the prediction error, which is the input from the inverse transform unit 206, and the prediction sample, which is the input from the prediction control unit 220. Also, the addition unit 208 outputs the reconstructed block to the block memory 210 and the loop filter unit 212.
[0223] [Block memory]
[0224] The block memory 210 is a storage unit for storing blocks within the decoded picture (hereinafter referred to as the current picture) that are referenced in intra - prediction. Specifically, the block memory 210 stores the reconstructed blocks output from the addition unit 208.
[0225] [Loop filter unit]
[0226] The loop filter unit 212 applies loop filtering to the block reconstructed by the addition unit 208 and outputs the filtered reconstructed block to the frame memory 214, the display device, etc.
[0227] When the information decoded from the coded bitstream indicating the on / off of ALF indicates that ALF is on, one filter is selected from a plurality of filters based on the direction and activity of the local gradient, and the selected filter is applied to the reconstructed block.
[0228] [Frame memory]
[0229] The frame memory 214 is a storage unit for storing reference pictures used in inter - prediction, and is sometimes referred to as a frame buffer. Specifically, the frame memory 214 stores the reconstructed blocks filtered by the loop filter unit 212.
[0230] [Intra - prediction unit]
[0231] The intra prediction unit 216 generates a prediction signal (intra prediction signal) by performing intra prediction with reference to the blocks within the current picture stored in the block memory 210, based on the intra prediction mode decoded from the encoded bitstream. Specifically, the intra prediction unit 216 generates an intra prediction signal by performing intra prediction with reference to the samples (e.g., luminance values, chrominance differences) of the blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control unit 220.
[0232] In addition, when the intra prediction mode that refers to the luminance block is selected for the intra prediction of the chrominance block, the intra prediction unit 216 may also predict the chrominance component of the current block based on the luminance component of the current block.
[0233] Furthermore, when the information decoded from the encoded bitstream indicates the adoption of PDPC, the intra prediction unit 216 corrects the pixel value after intra prediction based on the gradients of the reference pixels in the horizontal / vertical directions.
[0234] [Inter prediction unit]
[0235] The inter prediction unit 218 predicts the current block with reference to the reference pictures stored in the frame memory 214. The prediction is performed in units of the current block or sub-blocks (e.g., 4×4 blocks) within the current block. For example, the inter prediction unit 218 performs motion compensation using the motion information (e.g., motion vectors) decoded from the encoded bitstream, thereby generating an inter prediction signal for the current block or sub-block, and outputs the inter prediction signal to the prediction control unit 220.
[0236] In addition, when the information decoded from the encoded bitstream indicates the adoption of the OBMC mode, the inter prediction unit 218 generates an inter prediction signal using not only the motion information of the current block obtained through motion estimation but also the motion information of the adjacent blocks.
[0237] Furthermore, when the information decoded from the encoded bitstream indicates the adoption of the FRUC mode, the inter prediction unit 218 performs motion estimation according to the pattern matching method (bidirectional matching or template matching) decoded from the encoded stream, thereby deriving the motion information. And the inter prediction unit 218 performs motion compensation using the derived motion information.
[0238] In addition, when the BIO mode is adopted, the inter prediction unit 218 derives the motion vector based on a model assuming uniform linear motion. Furthermore, when the information decoded from the encoded bitstream indicates the adoption of the affine motion compensation prediction mode, the inter prediction unit 218 derives the motion vector in units of sub-blocks based on the motion vectors of multiple adjacent blocks.
[0239] [Prediction control unit]
[0240] The prediction control unit 220 selects one of an intra prediction signal and an inter prediction signal, and outputs the selected signal as a prediction signal to the addition unit 208.
[0241] (Embodiment 2)
[0242] Regarding the encoding process and decoding process related to Embodiment 2, refer to Figure 11 and Figure 12 Specifically, regarding the encoding device and decoding device related to Embodiment 2, refer to Figure 15 and Figure 16 Specifically described.
[0243] [Encoding Process]
[0244] Figure 11 Represents the video encoding process related to Embodiment 2.
[0245] First, in step S1001, a first parameter for identifying a partitioning mode for partitioning a first block into a plurality of sub-blocks from among a plurality of partitioning modes is written into the bitstream. If a partitioning mode is used, the block is partitioned into a plurality of sub-blocks. If different partitioning modes are used, the block is partitioned into a plurality of sub-blocks having different shapes, different heights, or different widths.
[0246] Figure 28 Represents an example of a partitioning mode for partitioning an N×N pixel block in Embodiment 2. In Figure 28 , (a) to (h) represent mutually different partitioning modes. As Figure 28As shown, if the block mode (a) is used, a block of N×N pixels (for example, 16×16 pixels, and "N" can take any value that is an integer multiple of 4 from 8 to 128) is divided into two sub - blocks of N / 2×N pixels (for example, 8×16 pixels). If the block mode (b) is used, a block of N×N pixels is divided into a sub - block of N / 4×N pixels (for example, 4×16 pixels) and a sub - block of 3N / 4×N pixels (for example, 12×16 pixels). If the block mode (c) is used, a block of N×N pixels is divided into a sub - block of 3N / 4×N pixels (for example, 12×16 pixels) and a sub - block of N / 4×N pixels (for example, 4×16 pixels). If the block mode (d) is used, a block of N×N pixels is divided into sub - blocks of (N / 4)×N pixels (for example, 4×16 pixels), N / 2×N pixels (for example, 8×16 pixels), and N / 4×N pixels (for example, 4×16 pixels). If the block mode (e) is used, a block of N×N pixels is divided into two sub - blocks of N×N / 2 pixels (for example, 16×8 pixels). If the block mode (f) is used, a block of N×N pixels is divided into a sub - block of N×N / 4 pixels (for example, 16×4 pixels) and a sub - block of N×3N / 4 pixels (for example, 16×12 pixels). If the block mode (g) is used, a block of N×N pixels is divided into a sub - block of N×3N / 4 pixels (for example, 16×12 pixels) and a sub - block of N×N / 4 pixels (for example, 16×4 pixels). If the block mode (h) is used, a block of N×N pixels is divided into sub - blocks of N×N / 4 pixels (for example, 16×4 pixels), N×N / 2 pixels (for example, 16×8 pixels), and N×N / 4 pixels (for example, 16×4 pixels).
[0247] Next, in step S1002, it is judged whether the first parameter identifies the first block mode.
[0248] Next, in step S1003, based at least on the judgment of whether the first parameter identifies the first block mode, it is judged whether the second block mode is not selected as a candidate for dividing the second block.
[0249] Two different block - mode sets may divide a block into sub - blocks of the same shape and size. For example, as Figure 31A shown, the sub - blocks of (1b) and (2c) have the same shape and size. One block - mode set can contain at least two block modes. For example, as Figure 31A shown in (1a) and (1b) of Figure 31AAs shown in (2a), (2b), and (2c), other block pattern sets can follow the vertical binary tree splitting that includes vertical binary tree splitting of two sub-blocks. Each block pattern set results in sub-blocks of the same shape and size.
[0250] When selecting between two block pattern sets that divide a block into sub-blocks of the same shape and size and that are different binary numbers or different numbers of bits when encoded in a bitstream, select the block pattern set with fewer binary numbers or fewer bits. Additionally, the binary numbers and the number of bits correspond to the code amount.
[0251] When selecting between two block pattern sets that divide a block into sub-blocks of the same shape and size and that are the same binary number or the same number of bits when encoded in a bitstream, select the block pattern set that appears first in a specified order of multiple block pattern sets. The specified order can be, for example, an order based on the number of block patterns within each block pattern set.
[0252] Figure 31A and Figure 31B is a diagram showing an example of dividing a block into sub-blocks using a block pattern set with fewer binary numbers in the encoding of the block pattern. In this example, when the left N×N pixel block is vertically divided into two sub-blocks, the second block pattern for the right N×N pixel block is not selected in step (2c). This is because, in Figure 31B the encoding method of the block pattern, the second block pattern set (2a, 2b, 2c) requires more binary numbers for encoding the block pattern compared to the first block pattern set (1a, 1b).
[0253] Figures 32A to 32C is a diagram showing an example of dividing a block into sub-blocks using the block pattern set that appears first in a specified order of multiple block pattern sets. In this example, when the 2N×N / 2 pixel block is vertically divided into three sub-blocks, the second block pattern for the lower 2N×N / 2 pixel block is not selected in step (2c). This is because, in Figure 32B the encoding method of the block pattern, the second block pattern set (2a, 2b, 2c) has the same binary number as the first block pattern set (1a, 1b, 1c, 1d), and in Figure 32C the specified order of the block pattern sets shown, it appears after the first block pattern set (1a, 1b, 1c, 1d). The specified order of multiple block pattern sets can also be fixed and can be signaled within the bitstream.
[0254] Figure 20This shows an example in Embodiment 2 where the second block division mode is not selected for the division of a 2N×N pixel block as shown in step (2c). As Figure 20 shown, it is possible to use the first division method (i) to equally divide a 2N×2N pixel block (e.g., 16×16 pixels) into 4 sub-blocks of N×N pixels (e.g., 8×8 pixels) as in step (1a). Also, it is possible to use the second division method (ii) to horizontally equally divide a 2N×2N pixel block into 2 sub-blocks of 2N×N pixels (e.g., 16×8 pixels) as in step (2a). Here, in the second division method (ii), when the upper 2N×N pixel block (the first block) is vertically divided into 2 sub-blocks of N×N pixels by the first block division mode as in step (2b), in step (2c), the second block division mode for vertically dividing the lower 2N×N pixel block (the second block) into 2 sub-blocks of N×N pixels is not selected as a candidate for the possible block division mode. This is because sub-block sizes identical to those obtained by the four-fold division using the first division method (i) are generated.
[0255] As described above, in Figure 20 , when the first block is vertically equally divided into 2 sub-blocks if the first block division mode is used, and the second block adjacent to the first block in the vertical direction is vertically equally divided into 2 sub-blocks if the second block division mode is used, the second block division mode is not selected as a candidate.
[0256] Figure 21 This shows an example in Embodiment 2 where the second block division mode is not selected for the division of an N×2N pixel block as shown in step (2c). As Figure 21 shown, it is possible to use the first division method (i) to equally divide a 2N×2N pixel block into 4 sub-blocks of N×N pixels as in step (1a). Also, it is possible to use the second division method (ii) to vertically equally divide a 2N×2N pixel block into 2 sub-blocks of 2N×N pixels (e.g., 8×16 pixels) as in step (2a). In the second division method (ii), when the left N×2N pixel block (the first block) is horizontally divided into 2 sub-blocks of N×N pixels by the first block division mode as in step (2b), in step (2c), the second block division mode for horizontally dividing the right N×2N pixel block (the second block) into 2 sub-blocks of N×N pixels is not selected as a candidate for the possible block division mode. This is because sub-block sizes identical to those obtained by the four-fold division using the first division method (i) are generated.
[0257] As described above, in Figure 21In the case where, if the first block mode is used, the first block is equally divided into two sub-blocks in the horizontal direction, and if the second block mode is used, the second block adjacent to the first block in the horizontal direction is equally divided into two sub-blocks in the horizontal direction, the second block mode is not selected as a candidate.
[0258] Figure 40 Indicates an example of dividing a 4N×2N block in Figure 20 into three parts in a ratio of 1:2:1, such as N×2N, 2N×2N, N×2N. Here, when the upper block is divided into three parts, the block division mode of dividing the lower block into three parts in a ratio of 1:2:1 is not selected as a candidate for possible block division modes. The three-way division can also be in a ratio different from 1:2:1. Furthermore, it can be divided into more than three parts, it can also be divided into two parts, or it can be in a ratio different from 1:1, such as 1:2 or 1:3. Figure 40 This is an example of dividing first in the horizontal direction, but the same constraints can also be applied when dividing first in the vertical direction.
[0259] Figure 41 and Figure 42 Indicates an example of applying the same constraint when the first block is rectangular.
[0260] Figure 43 This is the second constraint example when a square is divided into three parts in the vertical direction and then equally divided into two parts in the horizontal direction. When applying Figure 43 the constraint, in Figure 40 , it is possible to select a block division mode of dividing the lower block of 4N×2N into three parts in a ratio of 1:2:1. It is also possible to separately encode information indicating which of the constraints of Figure 40 and Figure 43 is applied into the header information, etc. Alternatively, it is also possible to apply a constraint to reduce the code amount of the information representing the block division. For example, if it is assumed that the code amounts of the information representing the block division in Case 1 and Case 2 are as follows, the division in Case 1 is set to be valid and the division in Case 2 is set to be invalid. That is, apply Figure 43 the constraint.
[0261] (Case 1) (1) Divide the square into two parts in the horizontal direction, and then (2) divide the upper and lower two rectangular blocks vertically into three parts: (1) Direction information: 1 bit, division quantity information: 1 bit, (2) (Direction information: 1 bit, division quantity information: 1 bit)×2, a total of 6 bits
[0262] (Case 2) (1) Divide the square into two parts in the vertical direction, and then (2) divide the left, middle, and right rectangular blocks horizontally into two parts: (1) Direction information: 1 bit, division quantity information: 1 bit, (2) (Direction information: 1 bit, division quantity information: 1 bit)×3, a total of 8 bits
[0263] Alternatively, there is a case where the optimal block is determined while selecting the block mode in a prescribed order during encoding. For example, 2-way splitting may be attempted first, followed by 3-way or 4-way splitting (bisecting horizontally and vertically). At this time, before attempting 3-way splitting as in Figure 43 , the attempt starting from 2-way splitting as in Figure 40 has already been performed. Thus, in the attempt starting from 2-way splitting, the splitting of horizontally bisecting and then vertically trisecting the two upper and lower blocks has already been attempted, so the constraint of Figure 43 is applied. In this way, the constraint method for selection can also be determined based on a prescribed encoding method.
[0264] In Figure 44 , an example is shown where the block modes that can be selected in the second block mode for the same direction as the first block mode are restricted. Here, the first block mode is 3-way splitting in the vertical direction. At this time, 2-way splitting cannot be selected as the second block mode. On the other hand, for the vertical direction, which is a direction different from the first block mode, 2-way splitting can be selected ( Figure 45 ).
[0265] Figure 22 An example is shown where the second block mode is not selected for the splitting of an N×N pixel block as in step (2c) in Embodiment 2. As shown in Figure 22 , the first splitting method (i) can be used to vertically split a 2N×N pixel block (e.g., 16×8 pixels, and any value that is an integer multiple of 4 from 8 to 128 can be taken as the value of "N") into sub-blocks of N / 2×N pixels, N×N pixels, and N / 2×N pixels (e.g., sub-blocks of 4×8 pixels, 8×8 pixels, and 4×8 pixels). Also, the second splitting method (ii) can be used to split a 2N×N pixel block into two N×N pixel sub-blocks as in step (2a). In the first splitting method (i), the central N×N pixel block can be vertically split into two N / 2×N pixel (e.g., 4×8 pixel) sub-blocks in step (1b). In the second splitting method (ii), when the left N×N pixel block (the first block) is vertically split into two N / 2×N pixel sub-blocks as in step (2b), the block mode of vertically splitting the right N×N pixel block (the second block) into two N / 2×N pixel sub-blocks in step (2c) is not selected as a candidate for possible block modes. This is because sub-blocks of the same size as those obtained by the first splitting method (i) will be generated, i.e., four N / 2×N pixel sub-blocks.
[0266] As described above, in Figure 22In the case where, if the first block mode is used, the first block is equally divided into two sub-blocks in the vertical direction, and if the second block mode is used, the second block adjacent to the first block in the horizontal direction is equally divided into two sub-blocks in the vertical direction, the second block mode is not selected as a candidate.
[0267] Figure 23 This shows an example in Embodiment 2 where the second block mode is not selected for the division of an N×N pixel block as shown in step (2c). As Figure 23 shown, the first division method (i) can be used to divide an N×2N pixel (e.g., 8×16 pixels, and any value that is an integer multiple of 4 from 8 to 128 can be taken as the value of "N") into sub-blocks of N×N / 2 pixels, sub-blocks of N×N pixels, and sub-blocks of N×N / 2 pixels (e.g., sub-blocks of 8×4 pixels, sub-blocks of 8×8 pixels, and sub-blocks of 8×4 pixels). Also, the second division method can be used to divide it into two sub-blocks of N×N pixels as shown in step (2a). In the first division method (i), the central N×N pixel block can be divided into two sub-blocks of N×N / 2 pixels as shown in step (1b). In the second division method (ii), when the upper N×N pixel block (the first block) is horizontally divided into two sub-blocks of N×N / 2 pixels as shown in step (2b), in step (2c), the block division mode of horizontally dividing the lower N×N pixel block (the second block) into two sub-blocks of N×N / 2 pixels is not selected as a candidate for possible block division modes. This is because sub-blocks of the same size as those obtained by the first division method (i) will be generated, that is, four sub-blocks of N×N / 2 pixels.
[0268] As described above, in Figure 23 the case where, if the first block mode is used, the first block is equally divided into two sub-blocks in the horizontal direction, and if the second block mode is used, the second block adjacent to the first block in the vertical direction is equally divided into two sub-blocks in the horizontal direction, the second block mode is not selected as a candidate.
[0269] If it is determined that the second block mode is selected as a candidate for dividing the second block (No in S1003), then in step S1004, a block division mode is selected from among multiple block division modes that include the second block mode as a candidate. In step S1005, a second parameter indicating the selection result is written into the bitstream.
[0270] If it is determined that the second block mode is not selected as a candidate for dividing the second block (Yes in S1003), then in step S1006, a block division mode different from the second block mode is selected for dividing the second block. Among the block division modes selected here, the block is divided into sub-blocks having a different shape or different size compared to the sub-blocks generated by the second block mode.
[0271] Figure 24 This represents an example of using the selected block mode (when not selecting the second block mode as shown in step (3)) to divide a 2N×N pixel block in Embodiment 2. As Figure 24 shown, the selected block mode can divide the current 2N×N pixel block (the lower block in this example) as Figure 24 shown in (c) and (f) into 3 sub-blocks. The sizes of the 3 sub-blocks can be different. For example, among the 3 sub-blocks, the large sub-block can have a width / height that is 2 times that of the small sub-block. And for example, the selected block mode can also divide the current block as Figure 24 shown in (a), (b), (d), and (e) into 2 sub-blocks of different sizes (asymmetric binary tree). For example, when using an asymmetric binary tree, the large sub-block can have a width / height that is 3 times that of the small sub-block.
[0272] Figure 25 This represents an example of using the selected block mode (when not selecting the second block mode as shown in step (3)) to divide an N×2N pixel block in Embodiment 2. As Figure 25 shown, the selected block mode can divide the current N×2N pixel block (the right block in this example) as Figure 25 shown in (c) and (f) into 3 sub-blocks. The sizes of the 3 sub-blocks can be different. For example, among the 3 sub-blocks, the large sub-block can have a width / height that is 2 times that of the small sub-block. And for example, the selected block mode can also divide the current block as Figure 25 shown in (a), (b), (d), and (e) into 2 sub-blocks of different sizes (asymmetric binary tree). For example, when using an asymmetric binary tree, the large sub-block can have a width / height that is 3 times that of the small sub-block.
[0273] Figure 26 This represents an example of using the selected block mode (when not selecting the second block mode as shown in step (3)) to divide an N×N pixel block in Embodiment 2. As Figure 26 shown, in step (1), the 2N×N pixel block is vertically divided into 2 N×N pixel sub-blocks, and in step (2), the left N×N pixel block is vertically divided into 2 N / 2×N pixel sub-blocks. In step (3), the selected block mode for the current N×N pixel block (the left block in this example) can be used to divide the current block as Figure 26 shown in (c) and (f) into 3 sub-blocks. The sizes of the 3 sub-blocks can be different. For example, among the 3 sub-blocks, the large sub-block can have a width / height that is 2 times that of the small sub-block. And for example, the selected block mode can also divide the current block as Figure 26is divided into two sub-blocks (asymmetric binary tree) with different sizes as shown in (a), (b), (d) and (e). For example, in the case of using an asymmetric binary tree, the large sub-block can have three times the width / height of the small sub-block.
[0274] Figure 27 This represents an example of dividing an N×N pixel block using the selected block partitioning mode when not selecting the second block partitioning mode as shown in step (3) in Embodiment 2. As Figure 27 shown, in step (1), the N×2N pixel block is horizontally divided into two N×N pixel sub-blocks, and in step (2), the upper N×N pixel block is horizontally divided into two N×N / 2 pixel sub-blocks. In step (3), the selected block partitioning mode for the current N×N pixel block (the lower block in this example) can be used to divide the current block into three sub-blocks as shown in Figure 27 (c) and (f). The sizes of the three sub-blocks can be different. For example, among the three sub-blocks, the large sub-block can have twice the width / height of the small sub-block. And for example, the selected block partitioning mode can also divide the current block into two sub-blocks with different sizes (asymmetric binary tree) as shown in Figure 27 (a), (b), (d) and (e). For example, in the case of using an asymmetric binary tree, the large sub-block can have three times the width / height of the small sub-block.
[0275] Figure 17 This represents the possible positions of the first parameter in the compressed video stream. As Figure 17 shown, the first parameter can be configured in the video parameter set, sequence parameter set, picture parameter set, slice header or coding tree unit. The first parameter can represent the method of dividing a block into multiple sub-blocks. For example, the first parameter can include a flag indicating whether to divide the block horizontally or vertically. The first parameter can also include a parameter indicating whether to divide the block into more than two sub-blocks.
[0276] Figure 18 This represents the possible positions of the second parameter in the compressed video stream. As Figure 18 shown, the second parameter can be configured in the video parameter set, sequence parameter set, picture parameter set, slice header or coding tree unit. The second parameter can represent the method of dividing a block into multiple sub-blocks. For example, the second parameter can include a flag indicating whether to divide the block horizontally or vertically. The second parameter can also include a parameter indicating whether to divide the block into more than two sub-blocks. As Figure 19 shown, the second parameter follows the first parameter in the bitstream.
[0277] The first block and the second block are different blocks. The first block and the second block may be included in the same frame. For example, the first block may be an adjacent block above the second block. And for example, the first block may also be an adjacent block to the left of the second block.
[0278] In step S1007, the second block is divided into sub-blocks using the selected block partitioning mode. In step S1008, the divided blocks are encoded.
[0279] [Encoding device]
[0280] Figure 15 is a block diagram showing the structure of the video / image encoding device according to Embodiment 2 or 3.
[0281] The video encoding device 5000 is a device for encoding an input video / image for each block to generate an encoded output bitstream. As Figure 15 shown, the video encoding device 5000 includes a transform unit 5001, a quantization unit 5002, an inverse quantization unit 5003, an inverse transform unit 5004, a block memory 5005, a frame memory 5006, an intra prediction unit 5007, an inter prediction unit 5008, an entropy encoding unit 5009, and a block segmentation determination unit 5010.
[0282] The input video is input to an adder, and the added value is output to the transform unit 5001. The transform unit 5001 transforms the added value into frequency coefficients based on the block partitioning mode derived by the block segmentation determination unit 5010, and outputs the frequency coefficients to the quantization unit 5002. The block partitioning mode can be associated with a block partitioning mode, a block partitioning type, or a block partitioning direction. The quantization unit 5002 quantizes the input quantization coefficients and outputs the quantized values to the inverse quantization unit 5003 and the entropy encoding unit 5009.
[0283] The inverse quantization unit 5003 inverse quantizes the quantized values output from the quantization unit 5002 and outputs the frequency coefficients to the inverse transform unit 5004. The inverse transform unit 5004 performs an inverse frequency transform on the frequency coefficients based on the block segmentation mode derived by the block segmentation determination unit 5010, transforms the frequency coefficients into sample values of the bitstream, and outputs the sample values to the adder.
[0284] The adder adds the sample values of the bitstream output from the inverse transform unit 5004 to the predicted video / image values output from the intra / inter prediction units 5007 and 5008, and outputs the added value to the block memory 5005 or the frame memory 5006 for further prediction. The block segmentation determination unit 5010 collects block information from the block memory 5005 or the frame memory 5006, and derives a block partitioning mode and parameters related to the block partitioning mode. If the derived block partitioning mode is used, the block is divided into a plurality of sub-blocks. The intra / inter prediction units 5007 and 5008 search among the video / images stored in the block memory 5005 or the video / images in the frame memory 5006 reconstructed by the block partitioning mode derived by the block segmentation determination unit 5010, and estimate, for example, the video / image region most similar to the input video / image to be predicted.
[0285] The entropy encoding unit 5009 encodes the quantization values output from the quantization unit 5002, encodes the parameters from the block segmentation determination unit 5010, and outputs a bitstream.
[0286] [Decoding process]
[0287] Figure 12 Indicates the video decoding process related to Embodiment 2.
[0288] First, in step S2001, a first parameter is interpreted from the bitstream. The first parameter identifies a partitioning mode for dividing a first block into sub-blocks from among a plurality of partitioning modes. If the partitioning mode is used, the block is divided into sub-blocks, and if a different partitioning mode is used, the block is divided into sub-blocks having different shapes, different heights, or different widths.
[0289] Figure 28 Indicates an example of the partitioning mode for dividing an N×N pixel block in Embodiment 2. In Figure 28 ,(a) to (h) represent different partitioning modes. As Figure 28As shown, if the block mode (a) is used, a block of N×N pixels (for example, 16×16 pixels, and as the value of "N", any value that is an integer multiple of 4 from 8 to 128 can be taken) is divided into two sub - blocks of N / 2×N pixels (for example, 8×16 pixels). If the block mode (b) is used, a block of N×N pixels is divided into a sub - block of N / 4×N pixels (for example, 4×16 pixels) and a sub - block of 3N / 4×N pixels (for example, 12×16 pixels). If the block mode (c) is used, a block of N×N pixels is divided into a sub - block of 3N / 4×N pixels (for example, 12×16 pixels) and a sub - block of N / 4×N pixels (for example, 4×16 pixels). If the block mode (d) is used, a block of N×N pixels is divided into a sub - block of (N / 4)×N pixels (for example, 4×16 pixels), a sub - block of N / 2×N pixels (for example, 8×16 pixels), and a sub - block of N / 4×N pixels (for example, 4×16 pixels). If the block mode (e) is used, a block of N×N pixels is divided into two sub - blocks of N×N / 2 pixels (for example, 16×8 pixels). If the block mode (f) is used, a block of N×N pixels is divided into a sub - block of N×N / 4 pixels (for example, 16×4 pixels) and a sub - block of N×3N / 4 pixels (for example, 16×12 pixels). If the block mode (g) is used, a block of N×N pixels is divided into a sub - block of N×3N / 4 pixels (for example, 16×12 pixels) and a sub - block of N×N / 4 pixels (for example, 16×4 pixels). If the block mode (h) is used, a block of N×N pixels is divided into a sub - block of N×N / 4 pixels (for example, 16×4 pixels), a sub - block of N×N / 2 pixels (for example, 16×8 pixels), and a sub - block of N×N / 4 pixels (for example, 16×4 pixels).
[0290] Next, in step S2002, it is determined whether the first parameter identifies the first block mode.
[0291] Next, in step S2003, based at least on the determination of whether the first parameter identifies the first block mode, it is determined whether the second block mode is not selected as a candidate for dividing the second block.
[0292] Two different sets of block modes may divide a block into sub - blocks of the same shape and size. For example, as Figure 31A shown, the sub - blocks of (1b) and (2c) have the same shape and size. One set of block modes can contain at least two block modes. For example, as Figure 31A shown in (1a) and (1b) of, one set of block modes can then, following a vertical trinary - tree split, contain a vertical binary - tree split of the central sub - block and a non - split of the other sub - blocks. And for example, as Figure 31AAs shown in (2a), (2b), and (2c), other block pattern sets can then include a binary tree vertical split that divides into two sub-blocks by vertically splitting a binary tree. Any block pattern set results in sub-blocks of the same shape and size.
[0293] When selecting between two block pattern sets that divide a block into sub-blocks of the same shape and size and that are different binary numbers or different numbers of bits when encoded in a bitstream, select the block pattern set with fewer binary numbers or fewer bits.
[0294] When selecting between two block pattern sets that divide a block into sub-blocks of the same shape and size and that have the same number of bits or the same number of bits when encoded in a bitstream, select the block pattern set that appears first in a specified order of multiple block pattern sets. The specified order can be, for example, an order based on the number of block patterns within each block pattern set.
[0295] Figure 31A and Figure 31B is a diagram showing an example of dividing a block into sub-blocks using a block pattern set with fewer binary numbers in the encoding of the block pattern. In this example, when the left N×N pixel block is vertically divided into 2 sub-blocks, the second block pattern for the right N×N pixel block is not selected in step (2c). This is because, in Figure 31B the encoding method of the block pattern, the second block pattern set (2a, 2b, 2c) requires more binary numbers for encoding the block pattern compared to the first block pattern set (1a, 1b).
[0296] Figure 32A is a diagram showing an example of dividing a block into sub-blocks using the block pattern set that appears first in a specified order of multiple block pattern sets. In this example, when the 2N×N / 2 pixel block is vertically divided into 3 sub-blocks, the second block pattern for the lower 2N×N / 2 pixel block is not selected in step (2c). This is because, in Figure 32B the encoding method of the block pattern, the second block pattern set (2a, 2b, 2c) has the same binary number as the first block pattern set (1a, 1b, 1c, 1d), and in Figure 32C the specified order of the block pattern sets shown, it appears after the first block pattern set (1a, 1b, 1c, 1d). The specified order of multiple block pattern sets can also be fixed and can be signaled within the bitstream.
[0297] Figure 20 shows an example of not selecting the second block pattern for the division of the 2N×N pixel block as shown in step (2c) in Embodiment 2. As Figure 20As shown, it is possible to use the first splitting method (i) to equally split a block of 2N×2N pixels (e.g., 16×16 pixels) into 4 sub-blocks of N×N pixels (e.g., 8×8 pixels) as in step (1a). Also, it is possible to use the second splitting method (ii) to horizontally equally split a block of 2N×2N pixels into 2 sub-blocks of 2N×N pixels (e.g., 16×8 pixels) as in step (2a). In the second splitting method (ii), when the upper 2N×N pixel block (the first block) is vertically split into 2 sub-blocks of N×N pixels by the first block splitting pattern as in step (2b), the second block splitting pattern for vertically splitting the lower 2N×N pixel block (the second block) into 2 sub-blocks of N×N pixels is not selected as a candidate for the possible block splitting pattern in step (2c). This is because sub-blocks of the same size as those obtained by the four-way split using the first splitting method (i) are generated.
[0298] As described above, Figure 20 in a case where if the first block splitting pattern is used, the first block is equally split into 2 sub-blocks in the vertical direction, and if the second block splitting pattern is used, the second block adjacent to the first block in the vertical direction is equally split into 2 sub-blocks in the vertical direction, the second block splitting pattern is not selected as a candidate.
[0299] Figure 21 This shows an example in Embodiment 2 where the second block splitting pattern is not selected for splitting a block of N×2N pixels as in step (2c). As Figure 21 shown, it is possible to use the first splitting method (i) to equally split a block of 2N×2N pixels into 4 sub-blocks of N×N pixels as in step (1a). Also, it is possible to use the second splitting method (ii) to vertically equally split a block of 2N×2N pixels into 2 sub-blocks of 2N×N pixels (e.g., 8×16 pixels) as in step (2a). In the second splitting method (ii), when the left N×2N pixel block (the first block) is horizontally split into 2 sub-blocks of N×N pixels by the first block splitting pattern as in step (2b), the second block splitting pattern for horizontally splitting the right N×2N pixel block (the second block) into 2 sub-blocks of N×N pixels is not selected as a candidate for the possible block splitting pattern in step (2c). This is because sub-blocks of the same size as those obtained by the four-way split using the first splitting method (i) are generated.
[0300] As described above, Figure 21 in a case where if the first block splitting pattern is used, the first block is equally split into 2 sub-blocks in the horizontal direction, and if the second block splitting pattern is used, the second block adjacent to the first block in the horizontal direction is equally split into 2 sub-blocks in the horizontal direction, the second block splitting pattern is not selected as a candidate.
[0301] Figure 22 This shows an example where the second block division mode is not selected for the division of an N×N pixel block as shown in step (2c) in Embodiment 2. As Figure 22 shown, the first division method (i) can be used to vertically divide a 2N×N pixel block (for example, 16×8 pixels, and any value that is an integer multiple of 4 from 8 to 128 can be taken as the value of "N") into sub-blocks of N / 2×N pixels, N×N pixels, and N / 2×N pixels (for example, sub-blocks of 4×8 pixels, 8×8 pixels, and 4×8 pixels). Also, the second division method (ii) can be used to divide a 2N×N pixel block into two N×N pixel sub-blocks as shown in step (2a). In the first division method (i), the central N×N pixel block can be vertically divided into two N / 2×N pixel (for example, 4×8 pixel) sub-blocks in step (1b). In the second division method (ii), when the left N×N pixel block (the first block) is vertically divided into two N / 2×N pixel sub-blocks as shown in step (2b), the block division mode of vertically dividing the right N×N pixel block (the second block) into two N / 2×N pixel sub-blocks in step (2c) is not selected as a candidate for the possible block division modes. This is because sub-block sizes identical to those obtained by the first division method (i) will be generated, that is, four N / 2×N pixel sub-blocks.
[0302] As described above, Figure 22 in a case where if the first block division mode is used, the first block is equally divided into two sub-blocks in the vertical direction, and if the second block division mode is used, the second block adjacent to the first block in the horizontal direction is equally divided into two sub-blocks in the vertical direction, the second block division mode is not selected as a candidate.
[0303] Figure 23 This shows an example where the second block division mode is not selected for the division of an N×N pixel block as shown in step (2c) in Embodiment 2. As Figure 23As shown, it is possible to use the first segmentation method (i) to divide an N×2N pixel (for example, 8×16 pixels, and any value that is an integer multiple of 4 from 8 to 128 can be taken as the value of "N") into sub-blocks of N×N / 2 pixels, sub-blocks of N×N pixels, and sub-blocks of N×N / 2 pixels (for example, sub-blocks of 8×4 pixels, sub-blocks of 8×8 pixels, and sub-blocks of 8×4 pixels) as in step (1a). Also, it is possible to use the second segmentation method to divide it into two sub-blocks of N×N pixels as in step (2a). In the first segmentation method (i), it is possible to divide the central block of N×N pixels into two sub-blocks of N×N / 2 pixels as in step (1b). In the second segmentation method (ii), when the upper N×N pixel block (the first block) is horizontally divided into two sub-blocks of N×N / 2 pixels as in step (2b), the block division mode in which the lower N×N pixel block (the second block) is horizontally divided into two sub-blocks of N×N / 2 pixels in step (2c) is not selected as a candidate for the possible block division mode. This is because sub-blocks of the same size as those obtained by the first segmentation method (i) will be generated, that is, four sub-blocks of N×N / 2 pixels.
[0304] As described above, Figure 23 in the case where, if the first block division mode is used, the first block is equally divided into two sub-blocks in the horizontal direction, and if the second block division mode is used, the second block adjacent to the first block in the vertical direction is equally divided into two sub-blocks in the horizontal direction, the second block division mode is not selected as a candidate.
[0305] If it is determined that the second block division mode is selected as a candidate for dividing the second block (No in S2003), then in step S2004, the second parameter is decoded from the bitstream, and a block division mode is selected from among the multiple block division modes including the second block division mode as a candidate.
[0306] If it is determined that the second block division mode is not selected as a candidate for dividing the second block (Yes in S2003), then in step S2005, a block division mode different from the second block division mode is selected to divide the second block. The block division mode selected here divides the block into sub-blocks having a different shape or different size compared to the sub-blocks generated by the second block division mode.
[0307] Figure 24 An example of using the block division mode selected when the second block division mode is not selected to divide a 2N×N pixel block as in step (3) in Embodiment 2 is shown. As Figure 24 shown, the selected block division mode can divide the current 2N×N pixel block (the lower block in this example) as Figure 24The division shown in (c) and (f) is into 3 sub - blocks. The sizes of the 3 sub - blocks can be different. For example, among the 3 sub - blocks, the large sub - block can have a width / height that is 2 times that of the small sub - block. And for example, the selected block - partitioning pattern can also divide the current block as shown in Figure 24 (a), (b), (d), and (e) into 2 sub - blocks of different sizes (asymmetric binary tree). For example, when using an asymmetric binary tree, the large sub - block can have a width / height that is 3 times that of the small sub - block.
[0308] Figure 25 This shows an example of using the selected block - partitioning pattern to divide a block of N×2N pixels without selecting the second block - partitioning pattern as shown in step (3) in Embodiment 2. As shown in Figure 25 The selected block - partitioning pattern can divide the current block of N×2N pixels (the right - hand block in this example) as shown in Figure 25 (c) and (f) into 3 sub - blocks. The sizes of the 3 sub - blocks can be different. For example, among the 3 sub - blocks, the large sub - block can have a width / height that is 2 times that of the small sub - block. And for example, the selected block - partitioning pattern can also divide the current block as shown in Figure 25 (a), (b), (d), and (e) into 2 sub - blocks of different sizes (asymmetric binary tree). For example, when using an asymmetric binary tree, the large sub - block can have a width / height that is 3 times that of the small sub - block.
[0309] Figure 26 This shows an example of using the selected block - partitioning pattern to divide a block of N×N pixels without selecting the second block - partitioning pattern as shown in step (3) in Embodiment 2. As shown in Figure 26 In step (1), the block of 2N×N pixels is vertically divided into 2 sub - blocks of N×N pixels. In step (2), the left - hand block of N×N pixels is vertically divided into 2 sub - blocks of N / 2×N pixels. In step (3), the selected block - partitioning pattern for the current block of N×N pixels (the left - hand block in this example) can be used to divide the current block as shown in Figure 26 (c) and (f) into 3 sub - blocks. The sizes of the 3 sub - blocks can be different. For example, among the 3 sub - blocks, the large sub - block can have a width / height that is 2 times that of the small sub - block. And for example, the selected block - partitioning pattern can also divide the current block as shown in Figure 26 (a), (b), (d), and (e) into 2 sub - blocks of different sizes (asymmetric binary tree). For example, when using an asymmetric binary tree, the large sub - block can have a width / height that is 3 times that of the small sub - block.
[0310] Figure 27 This shows an example of using the selected block - partitioning pattern to divide a block of N×N pixels without selecting the second block - partitioning pattern as shown in step (3) in Embodiment 2. As shown in Figure 27As shown, in step (1), a block of N×2N pixels is horizontally divided into two sub-blocks of N×N pixels. In step (2), the upper block of N×N pixels is horizontally divided into two sub-blocks of N×N / 2 pixels. In step (3), the selected partitioning mode for the current block of N×N pixels (the lower block in this example) can be used to divide the current block into three sub-blocks as shown in (c) and (f) of Figure 27 . The sizes of the three sub-blocks can be different. For example, among the three sub-blocks, the large sub-block has twice the width / height of the small sub-block. And for example, the selected partitioning mode can also divide the current block into two sub-blocks of different sizes (asymmetric binary tree) as shown in (a), (b), (d), and (e) of Figure 27 . For example, in the case of using an asymmetric binary tree, the large sub-block can have three times the width / height of the small sub-block.
[0311] Figure 17 Indicates the possible positions of the first parameter within the compressed video stream. As shown in Figure 17 , the first parameter can be configured in the video parameter set, sequence parameter set, picture parameter set, slice header, or coding tree unit. The first parameter can represent the method of dividing a block into multiple sub-blocks. For example, the first parameter can include a flag indicating whether to divide the block horizontally or vertically. The first parameter can also include a parameter indicating whether to divide the block into more than two sub-blocks.
[0312] Figure 18 Indicates the possible positions of the second parameter within the compressed video stream. As shown in Figure 18 , the second parameter can be configured in the video parameter set, sequence parameter set, picture parameter set, slice header, or coding tree unit. The second parameter can represent the method of dividing a block into multiple sub-blocks. For example, the second parameter can include a flag indicating whether to divide the block horizontally or vertically. The second parameter can also include a parameter indicating whether to divide the block into more than two sub-blocks. As shown in Figure 19 , the second parameter is configured in the bitstream following the first parameter.
[0313] The first block and the second block are different blocks. The first block and the second block can also be included in the same frame. For example, the first block can be a block adjacent above the second block. And for example, the first block can also be a block adjacent to the left of the second block.
[0314] In step S2006, the second block is divided into sub-blocks using the selected partitioning mode. In step S2007, the divided blocks are decoded.
[0315] [Decoding device]
[0316] Figure 16It is a block diagram showing the structure of the video / image decoding apparatus according to Embodiment 2 or 3.
[0317] The video decoding apparatus 6000 is an apparatus that decodes an input encoded bitstream for each block and outputs a video / image. The video decoding apparatus 6000 is as Figure 16 shown, and includes an entropy decoding unit 6001, an inverse quantization unit 6002, an inverse transform unit 6003, a block memory 6004, a frame memory 6005, an intra prediction unit 6006, an inter prediction unit 6007, and a block segmentation determination unit 6008.
[0318] The input encoded bitstream is input to the entropy decoding unit 6001. After the input encoded bitstream is input to the entropy decoding unit 6001, the entropy decoding unit 6001 decodes the input encoded bitstream, outputs the parameters to the block segmentation determination unit 6008, and outputs the decoded values to the inverse quantization unit 6002.
[0319] The inverse quantization unit 6002 performs inverse quantization on the decoded values and outputs the frequency coefficients to the inverse transform unit 6003. The inverse transform unit 6003 performs an inverse frequency transform on the frequency coefficients based on the block partitioning pattern derived by the block segmentation determination unit 6008, transforms the frequency coefficients into sample values, and outputs the sample values to the adder. The block partitioning pattern can be associated with a block partitioning pattern, a block partitioning type, or a block partitioning direction. The adder adds the sample values to the predicted video / image values output from the intra / inter prediction units 6006 and 6007, outputs the added values to the display, and outputs the added values to the block memory 6004 or the frame memory 6005 for further prediction. The block segmentation determination unit 6008 collects block information from the block memory 6004 or the frame memory 6005, and uses the parameters decoded by the entropy decoding unit 6001 to derive the block partitioning pattern. If the derived block partitioning pattern is used, the block is divided into a plurality of sub-blocks. Further, the intra / inter prediction units 6006 and 6007 perform prediction on the video / image region of the block to be decoded based on the video / image stored in the block memory 6004 or the video / image in the frame memory 6005 reconstructed according to the block partitioning pattern derived by the block segmentation determination unit 6008.
[0320] (Embodiment 3)
[0321] Refer to Figure 13 and Figure 14 to specifically describe the encoding process and decoding process according to Embodiment 3. Refer to Figure 15 and Figure 16 to specifically describe the encoding apparatus and decoding apparatus according to Embodiment 3.
[0322] [Encoding Process]
[0323] Figure 13Represents the video encoding process related to Embodiment 3.
[0324] First, in step S3001, a first parameter is written to the bitstream, which identifies, from among a plurality of block types, the block type used to divide the first block into sub-blocks.
[0325] In the next step S3002, a second parameter indicating the block division direction is written to the bitstream. The second parameter is configured in the bitstream following the first parameter. The block type and the block division direction can form a block division pattern together. The divided block represents the number and division ratio of sub-blocks used to divide the block.
[0326] Figure 29 Shows an example of the block type and block division direction used to divide an N×N pixel block in Embodiment 3. In Figure 29 , (1), (2), (3), and (4) are different block types, (1a), (2a), (3a), and (4a) are block division patterns with different block types in the vertical division direction, and (1b), (2b), (3b), and (4b) are block division patterns with different block types in the horizontal division direction. As Figure 29 shown, when the division ratio is 1:1 and the N×N pixel block is divided in a symmetric binary tree (i.e., 2 sub-blocks) along the vertical direction, the N×N pixel block is divided using the block division pattern (1a). When the division ratio is 1:1 and the N×N pixel block is divided in a symmetric binary tree (i.e., 2 sub-blocks) along the horizontal direction, the N×N pixel block is divided using the block division pattern (1b). When the division ratio is 1:3 and the N×N pixel block is divided in an asymmetric binary tree (i.e., 2 sub-blocks) along the vertical direction, the N×N pixel block is divided using the block division pattern (2a). When the division ratio is 1:3 and the N×N pixel block is divided in an asymmetric binary tree (i.e., 2 sub-blocks) along the horizontal direction, the N×N pixel block is divided using the block division pattern (2b). When the division ratio is 3:1 and the N×N pixel block is divided in an asymmetric binary tree (i.e., 2 sub-blocks) along the vertical direction, the N×N pixel block is divided using the block division pattern (3a). When the division ratio is 3:1 and the N×N pixel block is divided in an asymmetric binary tree (i.e., 2 sub-blocks) along the horizontal direction, the N×N pixel block is divided using the block division pattern (3b). When the division ratio is 1:2:1 and the N×N pixel block is divided in a ternary tree (i.e., 3 sub-blocks) along the vertical direction, the N×N pixel block is divided using the block division pattern (4a). When the division ratio is 1:2:1 and the N×N pixel block is divided in a ternary tree (i.e., 3 sub-blocks) along the horizontal direction, the N×N pixel block is divided using the block division pattern (4b).
[0327] Figure 17 Shows the possible positions of the first parameter in the compressed video stream. As Figure 17As shown, the first parameter can be configured within a video parameter set, a sequence parameter set, a picture parameter set, a slice header, or a coding tree unit. The first parameter can represent a method for dividing a block into multiple sub-blocks. For example, the first parameter can include a flag indicating whether to divide the block horizontally or vertically. The first parameter can also include a parameter indicating whether to divide the block into more than two sub-blocks.
[0328] Figure 18 Indicates a position where the second parameter within the compressed video stream can be considered. As Figure 18 shown, the second parameter can be configured within a video parameter set, a sequence parameter set, a picture parameter set, a slice header, or a coding tree unit. The second parameter can represent a method for dividing a block into multiple sub-blocks. For example, the second parameter can include a flag indicating whether to divide the block horizontally or vertically. The second parameter can also include a parameter indicating whether to divide the block into more than two sub-blocks. As Figure 19 shown, the second parameter is configured in the bitstream following the first parameter.
[0329] Figure 30 Indicates the advantages of encoding the block type before the block direction compared to encoding the block direction before the block type. In this example, when the horizontal block direction is invalidated due to an unsupported size (16×2 pixels), there is no need to encode the block direction. In this example, the block direction is determined to be the vertical block direction, and the horizontal block direction is invalid. When encoding the block type before the block direction, compared to encoding the block direction before the block type, the code bits brought by the encoding of the block direction are suppressed.
[0330] In this way, it is also possible to determine whether a block can be divided horizontally and vertically based on pre-determined conditions for block divisibility or non-divisibility. Then, when it is determined that the block can be divided only in one of the horizontal and vertical directions, the writing of the block direction to the bitstream can also be skipped. Furthermore, when it is determined that the block is not divisible in both the horizontal and vertical directions, in addition to skipping the writing of the block direction to the bitstream, the writing of the block type to the bitstream can also be skipped.
[0331] The pre-determined conditions for block divisibility or non-divisibility are defined, for example, by size (number of pixels) or number of divisions. The conditions for block divisibility or non-divisibility can also be pre-defined in the standard specification. Also, the conditions for block divisibility or non-divisibility can be included in a video parameter set, a sequence parameter set, a picture parameter set, a slice header, or a coding tree unit. The conditions for block divisibility or non-divisibility can be fixed for all blocks, or can be dynamically switched according to block characteristics (such as luminance and chrominance blocks) or picture characteristics (such as I, P, B pictures), etc.
[0332] In step S3003, the block is divided into sub-blocks using the identified block type and the indicated block direction. In step S3004, the divided blocks are encoded.
[0333] [Encoding device]
[0334] Figure 15 is a block diagram showing the structure of the video / image encoding device according to Embodiment 2 or 3.
[0335] The video encoding device 5000 is a device for encoding an input video / image for each block and generating an encoded output bit stream. As Figure 15 shown, the video encoding device 5000 includes a transformation unit 5001, a quantization unit 5002, an inverse quantization unit 5003, an inverse transformation unit 5004, a block memory 5005, a frame memory 5006, an intra prediction unit 5007, an inter prediction unit 5008, an entropy encoding unit 5009, and a block division determination unit 5010.
[0336] The input video is input to an adder, and the added value is output to the transformation unit 5001. The transformation unit 5001 transforms the added value into frequency coefficients based on the block division type and direction derived by the block division determination unit 5010, and outputs the frequency coefficients to the quantization unit 5002. The block division type and direction can be associated with the block division mode, the block division type, or the block division direction. The quantization unit 5002 quantizes the input quantization coefficients and outputs the quantization values to the inverse quantization unit 5003 and the entropy encoding unit 5009.
[0337] The inverse quantization unit 5003 inverse-quantizes the quantization values output from the quantization unit 5002 and outputs the frequency coefficients to the inverse transformation unit 5004. The inverse transformation unit 5004 performs an inverse frequency transformation on the frequency coefficients based on the block division type and direction derived by the block division determination unit 5010, transforms the frequency coefficients into sample values of the bit stream, and outputs the sample values to the adder.
[0338] The adder adds the sample values of the bitstream output from the intra / inter-frame prediction units 5007 and 5008 to the predicted video / image values, and outputs the added value to the block memory 5005 or the frame memory 5006 for further prediction. The block segmentation determination unit 5010 collects block information from the block memory 5005 or the frame memory 5006, and derives the block segmentation type and direction, as well as the parameters related to the block segmentation type and direction. If the derived block segmentation type and direction are used, the block is divided into a plurality of sub-blocks. The intra / inter-frame prediction units 5007 and 5008 search among the video / images stored in the block memory 5005 or the video / images in the frame memory 5006 reconstructed by the block segmentation type and direction derived by the block segmentation determination unit 5010, and estimate, for example, the video / image region most similar to the input video / image to be predicted.
[0339] The entropy encoding unit 5009 encodes the quantization values output from the quantization unit 5002, and encodes the parameters from the block segmentation determination unit 5010, and outputs a bitstream.
[0340] [Decoding process]
[0341] Figure 14 Represents the video decoding process related to Embodiment 3.
[0342] First, in step S4001, the first parameter is read from the bitstream, and the first parameter identifies the segmentation type for dividing the first block into sub-blocks from among a plurality of segmentation types.
[0343] In the next step S4002, the second parameter indicating the segmentation direction is read from the bitstream. The second parameter follows the first parameter in the bitstream. The segmentation type may also form a segmentation pattern together with the segmentation direction. The segmentation type indicates the number and segmentation ratio of the sub-blocks for dividing the block.
[0344] Figure 29 Shows an example of the segmentation type and segmentation direction for dividing an N×N pixel block in Embodiment 3. In Figure 29 ,(1), (2), (3) and (4) are different segmentation types, (1a), (2a), (3a) and (4a) are segmentation patterns with different segmentation types in the vertical direction, and (1b), (2b), (3b) and (4b) are segmentation patterns with different segmentation types in the horizontal direction. As Figure 29As shown, when the block ratio is 1:1 and the N×N pixel block is divided along the vertical direction in a symmetric binary tree (i.e., 2 sub-blocks), the N×N pixel block is divided using the block partitioning mode (1a). When the block ratio is 1:1 and the N×N pixel block is divided along the horizontal direction in a symmetric binary tree (i.e., 2 sub-blocks), the N×N pixel block is divided using the block partitioning mode (1b). When the block ratio is 1:3 and the N×N pixel block is divided along the vertical direction in an asymmetric binary tree (i.e., 2 sub-blocks), the N×N pixel block is divided using the block partitioning mode (2a). When the block ratio is 1:3 and the N×N pixel block is divided along the horizontal direction in an asymmetric binary tree (i.e., 2 sub-blocks), the N×N pixel block is divided using the block partitioning mode (2b). When the block ratio is 3:1 and the N×N pixel block is divided along the vertical direction in an asymmetric binary tree (i.e., 2 sub-blocks), the N×N pixel block is divided using the block partitioning mode (3a). When the block ratio is 3:1 and the N×N pixel block is divided along the horizontal direction in an asymmetric binary tree (i.e., 2 sub-blocks), the N×N pixel block is divided using the block partitioning mode (3b). When the block ratio is 1:2:1 and the N×N pixel block is divided along the vertical direction in a ternary tree (i.e., 3 sub-blocks), the N×N pixel block is divided using the block partitioning mode (4a). When the block ratio is 1:2:1 and the N×N pixel block is divided along the horizontal direction in a ternary tree (i.e., 3 sub-blocks), the N×N pixel block is divided using the block partitioning mode (4b).
[0345] Figure 17 Represents the possible positions of the first parameter within the compressed video stream. As Figure 17 shown, the first parameter can be configured in the video parameter set, sequence parameter set, picture parameter set, slice header, or coding tree unit. The first parameter can represent a method for dividing a block into multiple sub-blocks. For example, the first parameter can include an identifier for the above-mentioned block partitioning types. For example, the first parameter can include a flag indicating whether to divide the block horizontally or vertically. The first parameter can also include a parameter indicating whether to divide the block into more than 2 sub-blocks.
[0346] Figure 18 Represents the possible positions of the second parameter within the compressed video stream. As Figure 18 shown, the second parameter can be configured in the video parameter set, sequence parameter set, picture parameter set, slice header, or coding tree unit. The second parameter can represent a method for dividing a block into multiple sub-blocks. For example, the second parameter can include a flag indicating whether to divide the block horizontally or vertically. That is, the second parameter can include a parameter indicating the block partitioning direction. The second parameter can also include a parameter indicating whether to divide the block into more than 2 sub-blocks. As Figure 19 shown, the second parameter is configured in the bitstream following the first parameter.
[0347] Figure 30Represents the advantages of encoding the block type before the block direction compared to the case of encoding the block direction before the block type. In this example, when the block direction in the horizontal direction is invalidated due to an unsupported size (16×2 pixels), there is no need to encode the block direction. In this example, the block direction is determined to be the block direction in the vertical direction, and the block direction in the horizontal direction is invalidated. Encoding the block type before the block direction suppresses the code bits resulting from the encoding of the block direction compared to the case of encoding the block direction before the block type.
[0348] In this way, it is also possible to determine whether a block can be divided in the horizontal and vertical directions respectively based on pre-determined conditions for block divisibility or indivisibility. Then, when it is determined that the block can be divided only in one of the horizontal and vertical directions, the reading of the block direction from the bitstream can also be skipped. Furthermore, when it is determined that the block cannot be divided in both the horizontal and vertical directions, the reading of the block type from the bitstream can also be skipped in addition to the reading of the block direction.
[0349] The pre-determined conditions for block divisibility or indivisibility are defined by, for example, the size (number of pixels) or the number of divisions. These conditions for block divisibility or indivisibility can also be pre-defined in the standard specifications. Also, the conditions for block divisibility or indivisibility can be included in the video parameter set, sequence parameter set, picture parameter set, slice header, or coding tree unit. The conditions for block divisibility or indivisibility can be fixed for all blocks, or can be dynamically switched according to the characteristics of the block (e.g., luminance and chrominance blocks) or the characteristics of the picture (e.g., I, P, B pictures), etc.
[0350] In step S4003, the block is divided into sub-blocks using the identified block type and the indicated division direction. In step S4004, the divided block is decoded.
[0351] [Decoding device]
[0352] Figure 16 Is a block diagram showing the structure of the video / image decoding device according to Embodiment 2 or 3.
[0353] The video decoding device 6000 is a device for decoding an input encoded bitstream for each block and outputting a video / image. The video decoding device 6000 is as Figure 16 shown and includes an entropy decoding unit 6001, an inverse quantization unit 6002, an inverse transform unit 6003, a block memory 6004, a frame memory 6005, an intra prediction unit 6006, an inter prediction unit 6007, and a block division determination unit 6008.
[0354] The input encoded bitstream is input to the entropy decoding unit 6001. After the input encoded bitstream is input to the entropy decoding unit 6001, the entropy decoding unit 6001 decodes the input encoded bitstream, outputs the parameters to the block segmentation determination unit 6008, and outputs the decoded values to the inverse quantization unit 6002.
[0355] The inverse quantization unit 6002 performs inverse quantization on the decoded values and outputs the frequency coefficients to the inverse transform unit 6003. Based on the block partitioning type and direction derived by the block segmentation determination unit 6008, the inverse transform unit 6003 performs an inverse frequency transform on the frequency coefficients, transforms the frequency coefficients into sample values, and outputs the sample values to the adder. The block partitioning type and direction can be associated with the block partitioning mode, block partitioning type, or block partitioning direction. The adder adds the sample values to the predicted video / image values output from the intra / inter prediction units 6006 and 6007, outputs the added values to the display, and outputs the added values to the block memory 6004 or the frame memory 6005 for further prediction. The block segmentation determination unit 6008 collects block information from the block memory 6004 or the frame memory 6005, and uses the parameters decoded by the entropy decoding unit 6001 to derive the block partitioning type and direction. If the derived block partitioning type and direction are used, the block is divided into multiple sub-blocks. Furthermore, the intra / inter prediction units 6006 and 6007 perform prediction on the video / image region of the block to be decoded according to the video / image stored in the block memory 6004 or the video / image in the frame memory 6005 reconstructed according to the block partitioning type and direction derived by the block segmentation determination unit 6008.
[0356] (Embodiment 4)
[0357] In the above embodiments, each functional block can generally be implemented by an MPU and a memory, etc. In addition, the processing of each functional block is generally implemented by a program execution unit such as a processor reading and executing software (program) recorded in a recording medium such as a ROM. This software can be distributed by downloading, etc., or can be recorded in a recording medium such as a semiconductor memory for distribution. In addition, of course, each functional block can also be implemented by hardware (special-purpose circuit).
[0358] In addition, the processing described in each embodiment can be implemented by centralized processing using a single device (system), or can also be implemented by distributed processing using multiple devices. In addition, the processor that executes the above program can be single or multiple. That is, both centralized processing and distributed processing can be performed.
[0359] The form of the present invention is not limited to the above embodiments, and various changes can be made, and they are also included in the scope of the form of the present invention.
[0360] Furthermore, application examples of the moving image encoding method (image encoding method) or moving image decoding method (image decoding method) shown in the above-described embodiments and a system using the same will be described. The system is characterized by including an image encoding device using the image encoding method, an image decoding device using the image decoding method, and an image encoding / decoding device having both. Regarding other configurations in the system, they can be appropriately changed according to circumstances.
[0361] [Usage Example]
[0362] Figure 33 FIG. is a diagram showing the overall configuration of a content supply system ex100 that realizes a content distribution service. The provision area of the communication service is divided into desired sizes, and base stations ex106, ex107, ex108, ex109, and ex110 that are fixed wireless stations are respectively provided in each unit.
[0363] In this content supply system ex100, various devices such as a computer ex111, a game machine ex112, a camera ex113, home appliances ex114, and a smart phone ex115 are connected via the Internet ex101 through an Internet service provider ex102 or a communication network ex104, and base stations ex106 to ex110. The content supply system ex100 may also connect by combining some of the above elements. Each device may also be directly or indirectly connected to each other via a telephone network or short-range wireless without passing through the base stations ex106 to ex110 that are fixed wireless stations. In addition, a streaming media server ex103 is connected to various devices such as a computer ex111, a game machine ex112, a camera ex113, home appliances ex114, and a smart phone ex115 via the Internet ex101 or the like. In addition, the streaming media server ex103 is connected to terminals in a hotspot in an aircraft ex117 via a satellite ex116.
[0364] In addition, a wireless access point or a hotspot or the like may be used instead of the base stations ex106 to ex110. In addition, the streaming media server ex103 may be directly connected to the communication network ex104 without passing through the Internet ex101 or the Internet service provider ex102, or may be directly connected to the aircraft ex117 without passing through the satellite ex116.
[0365] The camera ex113 is a device such as a digital camera that can perform still image photography and moving image photography. In addition, the smart phone ex115 is a smart phone, a portable phone, or a PHS (Personal Handyphone System) corresponding to a mobile communication system mode generally referred to as 2G, 3G, 3.9G, 4G, and 5G in the future.
[0366] The home appliance ex118 is a refrigerator or a device included in a household fuel cell cogeneration system, etc.
[0367] In the content supply system ex100, a terminal having a photographing function is connected to a streaming media server ex103 via a base station ex106 or the like, whereby live distribution or the like can be performed. In live distribution, terminals (such as a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, a smart phone ex115, and a terminal in an airplane ex117) perform the encoding process described in the above embodiments on still image or moving image content photographed by a user using the terminal, multiplex the video data obtained by encoding and the audio data obtained by encoding the sound corresponding to the video, and send the obtained data to the streaming media server ex103. That is, each terminal functions as an image encoding device according to one aspect of the present invention.
[0368] On the other hand, the streaming media server ex103 performs stream distribution on the content data sent by a requesting client. The client is a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, a smart phone ex115, or a terminal in an airplane ex117 that can decode the data after the above encoding process. Each device that receives the distributed data performs decoding processing on the received data and reproduces it. That is, each device functions as an image decoding device according to one aspect of the present invention.
[0369] [Decentralized processing]
[0370] In addition, the streaming media server ex103 may be a plurality of servers or a plurality of computers, and perform decentralized processing or recording and distribution of data. For example, the streaming media server ex103 may be implemented by a CDN (Contents Delivery Network), and content distribution is achieved through a network connecting many edge servers dispersed in the world to each other. In the CDN, a physically closer edge server is dynamically allocated according to the client. And by caching and distributing the content to the edge server, the delay can be reduced. In addition, in the case of a certain error or a change in the communication state due to an increase in traffic or the like, the processing can be decentralized using multiple edge servers, or the distribution main body can be switched to another edge server, or a part of the network with a fault can be bypassed and the distribution can be continued, so high-speed and stable distribution can be achieved.
[0371] In addition, not limited to distributed processing of its own, the encoding process of the captured data can be performed by each terminal, on the server side, or can be shared between them. As an example, usually two processing loops are performed in the encoding process. In the first loop, the complexity or code amount of the image in units of frames or scenes is detected. In addition, in the second loop, a process of improving the encoding efficiency while maintaining the image quality is performed. For example, by performing the first encoding process by the terminal and the second encoding process by the server that receives the content, it is possible to reduce the processing load in each terminal while improving the quality and efficiency of the content. In this case, if there is a request to receive and decode almost in real time, the data completed by the first encoding performed by the terminal can also be received and reproduced by other terminals, so more flexible real-time distribution can also be performed.
[0372] As other examples, the camera ex113 etc. extracts feature amounts from the image, compresses the data regarding the feature amounts as metadata, and sends it to the server. The server, for example, judges the importance of the target based on the feature amounts and switches the quantization accuracy etc., and performs compression corresponding to the meaning of the image. The feature amount data is particularly effective for improving the accuracy and efficiency of motion vector prediction during re-compression in the server. In addition, simple encoding such as VLC (Variable Length Coding) can be performed by the terminal, and encoding with a large processing load such as CABAC (Context Adaptive Binary Arithmetic Coding) can be performed by the server.
[0373] As other examples, in a stadium, shopping mall, factory, etc., there are cases where there are multiple video data obtained by multiple terminals capturing substantially the same scene. In this case, the multiple terminals that performed the shooting are used, and other terminals and servers that did not perform the shooting are used as needed. For example, distributed processing is performed by separately allocating the encoding process in units of GOP (Group of Picture), picture units, or tile units obtained by dividing the picture. As a result, it is possible to reduce the delay and better achieve real-time performance.
[0374] In addition, since the multiple video data are of substantially the same scene, the server can also manage and / or instruct to refer to the video data captured by each terminal with each other. Or, it can also be that the server receives the encoded data from each terminal and changes the reference relationship between the multiple data, or corrects or replaces the picture itself and re-encodes it. As a result, a stream with improved quality and efficiency of each data can be generated.
[0375] In addition, the server can also perform transcoding to change the encoding method of the video data and then distribute the video data. For example, the server can change the MPEG-like encoding method to the VP-like, or can change H.264 to H.265.
[0376] Thus, the encoding process can be performed by a terminal or one or more servers. Therefore, the following descriptions use terms such as "server" or "terminal" as the processing entity. However, part or all of the processing performed by the server can also be performed by the terminal, and part or all of the processing performed by the terminal can also be performed by the server. In addition, the same applies to the decoding process regarding these.
[0377] [3D, Multi-angle]
[0378] In recent years, the use of combining different scenes captured by multiple cameras ex113 and / or terminals such as smartphones ex115 that are roughly synchronized with each other, or images or videos of the same scene captured from different angles has increased. The videos captured by each terminal are combined based on the relative position relationship between the terminals obtained separately, or the regions where the feature points included in the videos are consistent, etc.
[0379] The server not only encodes two-dimensional moving images, but can also encode still images automatically or at a user-specified time based on scene analysis of the moving images and send them to the receiving terminal. When the server can obtain the relative position relationship between the shooting terminals, it can not only generate the three-dimensional shape of the scene based on the videos of the same scene captured from different angles in addition to two-dimensional moving images. In addition, the server can also encode the three-dimensional data generated by point cloud, etc. separately, and can also select or reconstruct from the videos captured by multiple terminals based on the results of identifying or tracking people or objects using the three-dimensional data to generate and send the video to the receiving terminal.
[0380] In this way, the user can not only arbitrarily select each image corresponding to each shooting terminal to view the scene, but also view the content of the video cut from an arbitrary viewpoint from the three-dimensional data reconstructed using multiple images or videos. Furthermore, similar to the video, sound can also be collected from multiple different angles, and the server multiplexes and sends the sound from a specific angle or space with the video in accordance with the video.
[0381] In addition, in recent years, content that establishes a correspondence between the real world and the virtual world such as Virtual Reality (VR) and Augmented Reality (AR) has also been popularized. In the case of VR images, the server separately produces viewpoint images for the right eye and the left eye, and can perform encoding that allows reference between the viewpoint videos through Multi-View Coding (MVC), etc., or can encode them as different streams without mutual reference. When decoding different streams, they can be reproduced synchronously according to the user's viewpoint to reproduce a virtual three-dimensional space.
[0382] In the case of an AR image, it is also possible that the server overlaps the virtual object information on the virtual space with the camera information of the real space based on the three-dimensional position or the movement of the user's viewpoint. The decoding device acquires or holds the virtual object information and the three-dimensional data, generates a two-dimensional image according to the movement of the user's viewpoint, and creates the overlapping data by smoothly connecting them. Alternatively, it is also possible that the decoding device sends the movement of the user's viewpoint to the server in addition to the delegation of the virtual object information, and the server creates the overlapping data according to the three-dimensional data held in the server, matches the received movement of the viewpoint, encodes the overlapping data, and distributes it to the decoding device. In addition, the overlapping data has an α value representing the transmittance in addition to RGB, and the server sets the α value of the part other than the target created according to the three-dimensional data to 0, etc., and encodes it in a state where it is transmitted through this part. Alternatively, the server can also set the RGB value of a specified value as the background like chroma keying, and generate data with the part other than the target set as the background color.
[0383] Similarly, the decoding process of the distributed data can be performed by each terminal as a client, on the server side, or they can be shared with each other. As an example, it is also possible that a certain terminal first sends a reception request to the server, and another terminal receives the content corresponding to the request and performs the decoding process, and sends the decoded signal to the device with a display. By dispersing the processing regardless of the performance of the communicable terminal itself and selecting appropriate content, it is possible to reproduce data with better image quality. In addition, as another example, it is also possible for a TV or the like to receive large-size image data, and for the personal terminal of the viewer to decode and display a part of the area such as tiles after the picture is segmented. Thus, while making the overall image shared, it is possible to confirm one's own responsible area or the area that one wants to confirm in more detail at hand.
[0384] In addition, it is envisioned that in the future, in a situation where multiple wireless communications of short, medium, or long distances can be used both indoors and outdoors, using a distribution system standard such as MPEG-DASH, while seamlessly receiving content while appropriately switching data for the connected communication. Thus, the user can not only use their own terminal, but also freely select decoding devices or display devices such as monitors installed indoors and outdoors to switch in real time. In addition, based on their own position information, etc., it is possible to switch the decoding terminal and the display terminal for decoding. Thus, it is also possible to display map information on a part of the wall or floor of a building next to a displayable device while moving towards the destination. In addition, based on the ease of access to the encoded data on the network, such as the encoded data being cached in a server that can be accessed from the receiving terminal in a short time, or the encoded data being replicated in an edge server of the content distribution service, it is possible to switch the bit rate of the received data.
[0385] [Scalable Coding]
[0386] Regarding the switching of content, use Figure 34 As shown, a scalable stream that is compression-encoded using the moving image encoding method described in each of the above embodiments will be described. For the server, there may be multiple streams with the same content but different qualities as separate streams, or it may be a structure that switches content by utilizing the characteristics of a temporally / spatially scalable stream achieved by hierarchical encoding as shown in the figure. That is, the decoding side can freely switch between low-resolution content and high-resolution content for decoding by determining which layer to decode based on internal factors such as performance and external factors such as the state of the communication bandwidth. For example, when wanting to view the subsequent video that was viewed on a smartphone ex115 while on the move on a device such as an Internet TV after returning home, the device only needs to decode the same stream to different layers, so the burden on the server side can be reduced.
[0387] Furthermore, in addition to the structure where pictures are encoded by each layer as described above to achieve the hierarchical nature where the enhancement layer exists above the base layer, it is also possible that the enhancement layer contains meta-information such as the statistical information of the image, and the decoding side generates high-quality content by super-resolution of the pictures in the base layer based on the meta-information. Super-resolution can be either an improvement in the signal-to-noise ratio at the same resolution or an expansion of the resolution. The meta-information includes information for determining linear or non-linear filter coefficients used in the super-resolution process, or information for determining parameter values in filter processing, machine learning, or least squares operations used in the super-resolution process, etc.
[0388] Or, it is also possible to divide the picture into tiles, etc. according to the meaning of objects, etc. within the image, and the decoding side only decodes a part of the area by selecting the tiles to be decoded. In addition, by saving the attributes of the object (person, car, ball, etc.) and the position within the image (coordinate position in the same image, etc.) as meta-information, the decoding side can determine the position of the desired object based on the meta-information and decide on the tiles including the object. For example, as Figure 35 shown, use a data storage structure different from the pixel data, such as the SEI message in HEVC, to store the meta-information. This meta-information represents, for example, the position, size, or color of the main object.
[0389] In addition, the meta-information can also be stored in units composed of multiple pictures, such as streams, sequences, or random access units. Thereby, the decoding side can obtain the moment when a specific person appears within the video, etc., and by matching with the information of the picture unit, can determine the pictures where the object exists and the position of the object within the pictures.
[0390] [Optimization of Web pages]
[0391] Figure 36 It is a diagram showing an example of a display screen of a web page in a computer ex111 or the like. Figure 37 It is a diagram showing an example of a display screen of a web page in a smart phone ex115 or the like. As Figure 36 and Figure 37 shown, there are cases where a web page includes a plurality of link images that are links to image content, and the visible manner thereof varies depending on the viewing device. When a plurality of link images can be seen on the screen, before the user explicitly selects a link image, or before the link image approaches near the center of the screen or the entire link image enters the screen, the display device (decoding device) displays the still image or I picture that each content has as a link image, or displays an image such as a gif animation using a plurality of still images or I pictures, or only receives the base layer and decodes and displays the image.
[0392] When a link image is selected by the user, the display device decodes the base layer with the highest priority. In addition, if there is information indicating that the content is scalable in the HTML constituting the web page, the display device may also decode to the enhancement layer. Furthermore, in order to ensure real-time performance or when the communication band is very tight before selection, the display device can reduce the delay between the decoding time and the display time of the first picture (the delay from the start of content decoding to the start of display) by only decoding and displaying the forward-referenced pictures (I pictures, P pictures, B pictures that only perform forward reference). In addition, the display device can also forcibly ignore the reference relationship of the pictures and set all B pictures and P pictures as forward references for rough decoding, and as the pictures received over time increase, perform normal decoding.
[0393] [Autonomous driving]
[0394] In addition, when receiving still images or video data such as two-dimensional or three-dimensional map information for autonomous driving or driving assistance of a vehicle, the receiving terminal can also receive information such as weather or construction information as meta information in addition to the image data belonging to one or more layers, and decode them in correspondence. In addition, the meta information can either belong to a layer or be multiplexed only with the image data.
[0395] In this case, since vehicles, drones, airplanes, etc. including the receiving terminal are moving, the receiving terminal can switch the base stations ex106 to ex110 to perform seamless reception and decoding by sending the location information of the receiving terminal at the time of the reception request. In addition, the receiving terminal can dynamically switch the degree to which the meta information is received or the degree to which the map information is updated according to the user's selection, the user's condition, or the state of the communication band.
[0396] As described above, in the content supply system ex100, the client can receive, decode, and reproduce the encoded information sent by the user in real time.
[0397] [Distribution of Personal Content]
[0398] In addition, in the content supply system ex100, not only high-quality, long-duration content provided by video distribution providers but also low-quality, short-duration content provided by individuals can be distributed via unicast or multicast. In addition, it is conceivable that such personal content will increase in the future. To make personal content better, the server can also perform encoding after editing. This can be achieved, for example, through the following structure.
[0399] During shooting in real time or after cumulative shooting, the server performs recognition processing such as shooting error, scene search, meaning analysis, and object detection based on the original image or encoded data. And based on the recognition results, the server manually or automatically corrects focus deviation or camera shake, deletes scenes with low importance such as scenes with lower brightness or out-of-focus compared to other pictures, emphasizes the edges of the object, or changes the color tone, etc. Based on the editing results, the server encodes the edited data. In addition, it is known that the viewing rate decreases if the shooting time is too long. The server can also automatically crop scenes with low importance as described above and scenes with little movement based on the image processing results according to the shooting time to make the content within a specific time range. Or, the server can also generate a summary based on the result of scene meaning analysis and encode it.
[0400] In addition, in the original state, personal content may be written with content that infringes copyright, the author's personality rights, or portrait rights, etc., and there may also be inconvenient situations for individuals such as the sharing range exceeding the desired range. Therefore, for example, the server can also encode by forcibly changing the faces of people in the peripheral part of the screen or at home to out-of-focus images. In addition, the server can also identify whether a face of a person different from the pre-registered person is captured in the image to be encoded, and in the case of capture, perform processing such as applying a mosaic to the face part. Or, as pre-processing or post-processing of encoding, from the perspective of copyright, etc., the user can specify the person or background area of the image that they want to process, and the server performs processing such as replacing the specified area with another image or blurring the focus. If it is a person, the image of the face part can be replaced while tracking the person in the moving image.
[0401] In addition, for the audiovisual of personal content with a small amount of data, real-time requirements are relatively high. Therefore, although it also depends on the bandwidth, the decoding device first receives, decodes, and reproduces the base layer with the highest priority. The decoding device can also receive the enhancement layer during this period. When the reproduction is looped and reproduced more than twice, for example, it reproduces a high-quality image including the enhancement layer. In this way, if it is a scalable-encoded stream, it can provide an experience where the moving image is rough at the stage when it is not selected or just started to be viewed, but the stream gradually becomes smooth and the image quality improves. In addition to scalable encoding, the same experience can also be provided when the first rough stream and the second stream encoded with reference to the first moving image are combined into one stream.
[0402] [Other usage examples]
[0403] In addition, these encoding or decoding processes are usually processed in the LSIex500 provided in each terminal. The LSIex500 can be either a single-chip or a multi-chip structure. Additionally, software for moving image encoding or decoding can also be loaded into a certain recording medium (such as CD-ROM, floppy disk, hard disk, etc.) that can be read by a computer ex111, etc., and the encoding process and decoding process can be performed using this software. Furthermore, when the smartphone ex115 is equipped with a camera, it is also possible to transmit the moving image data obtained by this camera. The moving image data at this time is the data after being encoded by the LSIex500 provided in the smartphone ex115.
[0404] In addition, the LSIex500 can also be structured to download and activate application software. In this case, the terminal first determines whether the terminal corresponds to the encoding method of the content or has the ability to execute a specific service. When the terminal does not correspond to the encoding method of the content or does not have the ability to execute a specific service, the terminal downloads the codec or application software, and then performs content acquisition and reproduction.
[0405] In addition, it is not limited to the content supply system ex100 via the Internet ex101. It is also possible to incorporate at least one of the moving image encoding device (image encoding device) or the moving image decoding device (image decoding device) of the above-described embodiments into a digital broadcast system. Since the broadcast radio wave uses a satellite, etc., to carry multiplexed data that multiplexes video and audio for transmission and reception, compared with the structure of the content supply system ex100 that is easy for unicast, there are differences suitable for multicast, but the same applications can be made for the encoding process and the decoding process.
[0406] [Hardware structure]
[0407] Figure 38 It is a diagram showing the smartphone ex115. In addition,Figure 39 This is a diagram showing the structural example of the smart phone ex115. The smart phone ex115 has an antenna ex450 for transmitting and receiving radio waves with the base station ex110, a camera unit ex465 capable of shooting images and still pictures, and a display unit ex458 for displaying the images shot by the camera unit ex465 and decoding the data such as the images received by the antenna ex450. The smart phone ex115 also includes an operation unit ex466 such as a touch panel, a sound output unit ex457 such as a speaker for outputting sound or audio, a sound input unit ex456 such as a microphone for inputting sound, a memory unit ex467 capable of storing the shot images or still pictures, the recorded sound, the received images or still pictures, the encoded or decoded data such as emails, or a slot unit ex464 as an interface unit with the SIM ex468, and the SIM ex468 is used to identify the user and perform authentication for accessing various data represented by the network. In addition, an external memory can be used instead of the memory unit ex467.
[0408] In addition, the main control unit ex460 that comprehensively controls the display unit ex458, the operation unit ex466, etc. is connected to the power supply circuit unit ex461, the operation input control unit ex462, the image signal processing unit ex455, the camera interface unit ex463, the display control unit ex459, the modulation / demodulation unit ex452, the multiplexing / demultiplexing unit ex453, the sound signal processing unit ex454, the slot unit ex464, and the memory unit ex467 via the bus ex470.
[0409] If the power key is turned on by the user's operation, the power supply circuit unit ex461 starts the smart phone ex115 to an operable state by supplying power to each unit from the battery pack.
[0410] The smart phone ex115 performs processing such as calls and data communication under the control of a main control unit ex460 having a CPU, ROM, RAM, etc. During a call, the voice signal collection unit ex456 collects voice signals, which are then converted into digital voice signals by the voice signal processing unit ex454. These digital voice signals are subjected to spread spectrum processing by the modulation / demodulation unit ex452. After digital-to-analog conversion processing and frequency conversion processing are performed by the transmission / reception unit ex451, the signals are transmitted via the antenna ex450. In addition, received data is amplified and subjected to frequency conversion processing and analog-to-digital conversion processing. The modulation / demodulation unit ex452 performs inverse spread spectrum processing, and the voice signal processing unit ex454 converts the signals into analog voice signals, which are then output from the voice output unit ex457. During data communication, operations on the operation unit ex466 of the main body, etc., are used to send text, still images, or video data to the main control unit ex460 via the operation input control unit ex462, and similar transmission and reception processing is performed. In the data communication mode, when sending video, still images, or video and sound, the video signal processing unit ex455 compresses and encodes the video signals stored in the memory unit ex467 or the video signals input from the camera unit ex465 using the moving image encoding method shown in the above embodiments, and sends the encoded video data to the multiplexing / demultiplexing unit ex453. In addition, the voice signal processing unit ex454 encodes the voice signals collected by the voice input unit ex456 during the process of shooting video, still images, etc. by the camera unit ex465, and sends the encoded voice data to the multiplexing / demultiplexing unit ex453. The multiplexing / demultiplexing unit ex453 multiplexes the encoded video data and the encoded voice data in a specified manner, and the modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmission / reception unit ex451 perform modulation processing and conversion processing, and the signals are transmitted via the antenna ex450.
[0411] When an image attached to an email or a chat tool, or an image linked on a web page, etc. is received, in order to decode the multiplexed data received via the antenna ex450, the multiplexing / demultiplexing unit ex453 separates the multiplexed data into a bit stream of video data and a bit stream of audio data by demultiplexing the multiplexed data, supplies the encoded video data to the video signal processing unit ex455 via the synchronization bus ex470, and supplies the encoded audio data to the audio signal processing unit ex454. The video signal processing unit ex455 decodes the video signal by a video decoding method corresponding to the moving image encoding method described in the above embodiments, and displays the video or still image included in the linked moving image file from the display unit ex458 via the display control unit ex459. In addition, the audio signal processing unit ex454 decodes the audio signal and outputs the audio from the audio output unit ex457. In addition, since real-time streaming media is becoming popular, depending on the user's situation, there may be occasions where the reproduction of audio is inappropriate in society. Therefore, as an initial value, a structure in which only the video data is reproduced without reproducing the audio signal is preferable. It is also possible to reproduce the audio synchronously only when the user performs an operation such as clicking on the video data.
[0412] In addition, here, the smart phone ex115 has been described as an example, but as a terminal, in addition to the transceiver terminal having both an encoder and a decoder, three installation forms can be considered: a transmitting terminal having only an encoder and a receiving terminal having only a decoder. Furthermore, in a digital broadcast system, it has been described that multiplexed data in which audio data, etc. is multiplexed in video data is received and transmitted, but in addition to audio data, character data related to the video, etc. can also be multiplexed in the multiplexed data, and it is also possible to receive or transmit the video data itself instead of the multiplexed data.
[0413] In addition, it has been described that the main control unit ex460 including the CPU controls the encoding or decoding process, but in many cases, the terminal has a GPU. Therefore, it is also possible to configure a structure in which the performance of the GPU is used to process a larger area together by a memory shared by the CPU and the GPU, or a memory that manages addresses in a shared manner. Thereby, the encoding time can be shortened, real-time performance can be ensured, and low latency can be achieved. In particular, it is more effective if the processes of motion estimation, deblocking filtering, SAO (Sample Adaptive Offset), and transform / quantization are performed together by the GPU in units of pictures instead of by the CPU.
[0414] The encoding device according to an embodiment of the present disclosure may also be an encoding device that encodes a picture, and includes a processor and a memory; the processor has: a block division determination unit that divides the picture read from the memory into a plurality of blocks by using a set of block division patterns obtained by combining one or more block division patterns, where the block division patterns define division types; and an encoding unit that encodes the plurality of blocks; the set of block division patterns is composed of a first block division pattern and a second block division pattern, the first block division pattern defines a division direction and a division number for dividing a first block, and the second block division pattern defines a division direction and a division number for dividing a second block, which is one of the blocks obtained after dividing the first block; in the block division determination unit, when the division number of the first block division pattern is 3, the second block is the central block among the blocks obtained after dividing the first block, and the division direction of the second block division pattern is the same as the division direction of the first block division pattern, the second block division pattern only includes a block division pattern with a division number of 3.
[0415] The parameter for identifying the second block division pattern in the encoding device according to an embodiment of the present disclosure may also include a first flag indicating in which direction, the horizontal direction or the vertical direction, the block is divided, and does not include a second flag indicating the division number for dividing the block.
[0416] The encoding device according to an embodiment of the present disclosure may also be an encoding device that encodes a picture, and includes a processor and a memory; the processor has: a block division determination unit that divides the picture read from the memory into a plurality of blocks by using a set of block division patterns obtained by combining one or more block division patterns, where the block division patterns define division types; and an encoding unit that encodes the plurality of blocks; the set of block division patterns is composed of a first block division pattern and a second block division pattern, the first block division pattern defines a division direction and a division number for dividing a first block, and the second block division pattern defines a division direction and a division number for dividing a second block, which is one of the blocks obtained after dividing the first block; when the division number of the first block division pattern is 3, the second block is the central block among the blocks obtained after dividing the first block, and the division direction of the second block division pattern is the same as the division direction of the first block division pattern, the block division determination unit does not use the second block division pattern with a division number of 2.
[0417] The encoding device according to an embodiment of the present disclosure may also be an encoding device that encodes an image, and includes a processor and a memory; the processor has: a block division determination unit that divides the image read from the memory into a plurality of blocks by using a set of block division patterns obtained by combining one or more block division patterns, where the block division patterns define division types; and an encoding unit that encodes the plurality of blocks; the set of block division patterns includes a first block division pattern and a second block division pattern that respectively define a division direction and a division quantity; the block division determination unit restricts the use of the second block division pattern with the division quantity of 2.
[0418] The parameter for identifying the second block division pattern in the encoding device according to an embodiment of the present disclosure may also include a first flag indicating in which direction, horizontal or vertical, the block is divided, and a second flag indicating whether the block is divided into two or more blocks.
[0419] The above parameter in the encoding device according to an embodiment of the present disclosure may also be configured in slice data.
[0420] The encoding device according to an embodiment of the present disclosure may also be an encoding device that encodes an image, and includes a processor and a memory; the processor has: a block division determination unit that divides the image read from the memory into a block set composed of a plurality of blocks by using a set of block division patterns obtained by combining one or more block division patterns, where the block division patterns define division types; and an encoding unit that encodes the plurality of blocks; when the first block set obtained by using the first set of block division patterns is the same as the second block set obtained by using the second set of block division patterns, the block division determination unit performs division only by using one of the first set of block division patterns and the second set of block division patterns.
[0421] The block division determination unit in the encoding device according to an embodiment of the present disclosure may also perform division by using the block division pattern set with the smaller one of the first code quantity and the second code quantity based on the first code quantity of the first set of block division patterns and the second code quantity of the second set of block division patterns.
[0422] The block division determination unit in the encoding device according to an embodiment of the present disclosure may also perform division by using the block division pattern set that appears first in a preset order in the first set of block division patterns and the second set of block division patterns when the first code quantity is equal to the second code quantity based on the first code quantity of the first set of block division patterns and the second code quantity of the second set of block division patterns.
[0423] The decoding device according to an embodiment of the present disclosure may also be a decoding device that decodes an encoded signal, and includes a processor and a memory; the processor has: a block division determination unit that divides the encoded signal read from the memory into a plurality of blocks by using a set of block division patterns obtained by combining one or more block division patterns, where the block division patterns define division types; and a decoding unit that decodes the plurality of blocks; the set of block division patterns is composed of a first block division pattern and a second block division pattern, the first block division pattern defines a division direction and a division number for dividing a first block, and the second block division pattern defines a division direction and a division number for dividing a second block that is one of the blocks obtained after dividing the first block; in the block division determination unit, when the division number of the first block division pattern is 3, the second block is the central block among the blocks obtained after dividing the first block, and the division direction of the second block division pattern is the same as the division direction of the first block division pattern, the second block division pattern only includes a block division pattern with a division number of 3.
[0424] The parameter for identifying the second block division pattern in the decoding device according to an embodiment of the present disclosure may also include a first flag indicating in which direction, the horizontal direction or the vertical direction, the block is divided, and does not include a second flag indicating the division number for dividing the block.
[0425] The decoding device according to an embodiment of the present disclosure may also be a decoding device that decodes an encoded signal, and includes a processor and a memory; the processor has: a block division determination unit that divides the encoded signal read from the memory into a plurality of blocks by using a set of block division patterns obtained by combining one or more block division patterns, where the block division patterns define division types; and a decoding unit that decodes the plurality of blocks; the set of block division patterns is composed of a first block division pattern and a second block division pattern, the first block division pattern defines a division direction and a division number for dividing a first block, and the second block division pattern defines a division direction and a division number for dividing a second block that is one of the blocks obtained after dividing the first block; when the division number of the first block division pattern is 3, the second block is the central block among the blocks obtained after dividing the first block, and the division direction of the second block division pattern is the same as the division direction of the first block division pattern, the block division determination unit does not use the second block division pattern with a division number of 2.
[0426] The decoding device according to an embodiment of the present disclosure may also be a decoding device that decodes an encoded signal, and includes a processor and a memory; the processor has: a block division determination unit that divides the encoded signal read from the memory into a plurality of blocks by using a set of block division patterns obtained by combining one or more block division patterns, where the block division patterns define division types; and a decoding unit that decodes the plurality of blocks; the set of block division patterns includes a first block division pattern and a second block division pattern that respectively define a division direction and a division number; the block division determination unit restricts the use of the second block division pattern with the division number of 2.
[0427] The parameter for identifying the second block division pattern in the decoding device according to an embodiment of the present disclosure may also include a first flag indicating in which direction, horizontal or vertical, the block is divided, and a second flag indicating whether the block is divided into two or more.
[0428] The above parameter in the decoding device according to an embodiment of the present disclosure may also be configured in the slice data.
[0429] The decoding device according to an embodiment of the present disclosure may also be a decoding device that decodes an encoded signal, and includes a processor and a memory; the processor has: a block division determination unit that divides the encoded signal read from the memory into a block set composed of a plurality of blocks by using a set of block division patterns obtained by combining one or more block division patterns, where the block division patterns define division types; and a decoding unit that decodes the plurality of blocks; when the first block set obtained by using the first set of block division patterns is the same as the second block set obtained by using the second set of block division patterns, the block division determination unit divides using only one of the first set of block division patterns and the second set of block division patterns.
[0430] The block division determination unit in the decoding device according to an embodiment of the present disclosure may also divide using the block division pattern set with the smaller one of the first code amount of the first set of block division patterns and the second code amount of the second set of block division patterns based on the first code amount of the first set of block division patterns and the second code amount of the second set of block division patterns.
[0431] The block division determination unit in the decoding device according to an embodiment of the present disclosure may also divide using the block division pattern set that appears first in a preset order in the first set of block division patterns and the second set of block division patterns when the first code amount is equal to the second code amount based on the first code amount of the first set of block division patterns and the second code amount of the second set of block division patterns.
[0432] The encoding method according to an embodiment of the present disclosure may also be to use a set of block splitting patterns obtained by combining one or more block splitting patterns to split a picture read from a memory into a plurality of blocks, where the block splitting pattern defines a splitting type; encode the plurality of blocks; the set of block splitting patterns is composed of a first block splitting pattern and a second block splitting pattern, the first block splitting pattern defines a splitting direction and a splitting number for splitting a first block, and the second block splitting pattern defines a splitting direction and a splitting number for splitting a second block, which is one of the blocks obtained after splitting the first block; in the splitting, when the splitting number of the first block splitting pattern is 3, the second block is the central block among the blocks obtained after splitting the first block, and the splitting direction of the second block splitting pattern is the same as the splitting direction of the first block splitting pattern, the second block splitting pattern only includes a block splitting pattern with a splitting number of 3.
[0433] The parameter for identifying the second block splitting pattern in the encoding method according to an embodiment of the present disclosure may also include a first flag indicating in which direction, the horizontal direction or the vertical direction, the block is split, and does not include a second flag indicating the splitting number for splitting the block.
[0434] The encoding method according to an embodiment of the present disclosure may also be to have: a step of using a set of block splitting patterns obtained by combining one or more block splitting patterns to split a picture read from a memory into a plurality of blocks, where the block splitting pattern defines a splitting type; and a step of encoding the plurality of blocks; the set of block splitting patterns is composed of a first block splitting pattern and a second block splitting pattern, the first block splitting pattern defines a splitting direction and a splitting number for splitting a first block, and the second block splitting pattern defines a splitting direction and a splitting number for splitting a second block, which is one of the blocks obtained after splitting the first block; in the step of performing the splitting, when the splitting number of the first block splitting pattern is 3, the second block is the central block among the blocks obtained after splitting the first block, and the splitting direction of the second block splitting pattern is the same as the splitting direction of the first block splitting pattern, the second block splitting pattern with a splitting number of 2 is not used.
[0435] The encoding method according to an embodiment of the present disclosure may also be to have: a step of using a set of block splitting patterns obtained by combining one or more block splitting patterns to split a picture read from a memory into a plurality of blocks, where the block splitting pattern defines a splitting type; and a step of encoding the plurality of blocks; the set of block splitting patterns includes a first block splitting pattern and a second block splitting pattern that respectively define a splitting direction and a splitting number; in the step of performing the splitting, the use of the second block splitting pattern with a splitting number of 2 is restricted.
[0436] In the encoding method according to an embodiment of the present disclosure, the parameters for identifying the second block segmentation pattern may also include a first flag indicating in which direction, horizontal or vertical, the block is to be segmented, and a second flag indicating whether the block is to be segmented into two or more blocks.
[0437] In the encoding method according to an embodiment of the present disclosure, the above parameters may also be configured in the slice data.
[0438] The encoding method according to an embodiment of the present disclosure may also include: a step of using a set of block segmentation patterns obtained by combining one or more block segmentation patterns to divide a picture read from a memory into a set of blocks composed of a plurality of blocks, where the block segmentation pattern defines a segmentation type; and a step of encoding the plurality of blocks; in the step of performing the above division, when the first set of blocks obtained by using the first set of block segmentation patterns is the same as the second set of blocks obtained by using the second set of block segmentation patterns, only one of the first set of block segmentation patterns and the second set of block segmentation patterns is used for division.
[0439] In the step of performing the above division in the encoding method according to an embodiment of the present disclosure, division may also be performed using the set of block segmentation patterns with the smaller of the first code amount of the first set of block segmentation patterns and the second code amount of the second set of block segmentation patterns, based on the first code amount of the first set of block segmentation patterns and the second code amount of the second set of block segmentation patterns.
[0440] In the step of performing the above division in the encoding method according to an embodiment of the present disclosure, division may also be performed using the set of block segmentation patterns that appears first in a preset order among the first set of block segmentation patterns and the second set of block segmentation patterns when the first code amount is equal to the second code amount, based on the first code amount of the first set of block segmentation patterns and the second code amount of the second set of block segmentation patterns.
[0441] The decoding method according to an embodiment of the present disclosure may also be to use a set of block segmentation patterns obtained by combining one or more block segmentation patterns to divide an encoded signal read from a memory into a plurality of blocks, where the block segmentation pattern defines a segmentation type; decode the plurality of blocks; the set of block segmentation patterns is composed of a first block segmentation pattern and a second block segmentation pattern, the first block segmentation pattern defines a segmentation direction and a segmentation number for dividing a first block, and the second block segmentation pattern defines a segmentation direction and a segmentation number for dividing a second block, which is one of the blocks obtained after the segmentation of the first block; in the above segmentation, when the segmentation number of the first block segmentation pattern is 3, the second block is the central block among the blocks obtained after the segmentation of the first block, and the segmentation direction of the second block segmentation pattern is the same as the segmentation direction of the first block segmentation pattern, the second block segmentation pattern only includes a block segmentation pattern with a segmentation number of 3.
[0442] The parameter for identifying the second block segmentation pattern in the decoding method according to an embodiment of the present disclosure may also include a first flag indicating in which direction, the horizontal direction or the vertical direction, the block is segmented, and does not include a second flag indicating the segmentation number for segmenting the block.
[0443] The decoding method according to an embodiment of the present disclosure may also be to have: a step of using a set of block segmentation patterns obtained by combining one or more block segmentation patterns to divide an encoded signal read from a memory into a plurality of blocks, where the block segmentation pattern defines a segmentation type; and a step of decoding the plurality of blocks; the set of block segmentation patterns is composed of a first block segmentation pattern and a second block segmentation pattern, the first block segmentation pattern defines a segmentation direction and a segmentation number for dividing a first block, and the second block segmentation pattern defines a segmentation direction and a segmentation number for dividing a second block, which is one of the blocks obtained after the segmentation of the first block; in the step of performing the above segmentation, when the segmentation number of the first block segmentation pattern is 3, the second block is the central block among the blocks obtained after the segmentation of the first block, and the segmentation direction of the second block segmentation pattern is the same as the segmentation direction of the first block segmentation pattern, the second block segmentation pattern with a segmentation number of 2 is not used.
[0444] The decoding method according to an embodiment of the present disclosure may also be to have: a step of using a set of block segmentation patterns obtained by combining one or more block segmentation patterns to divide an encoded signal read from a memory into a plurality of blocks, where the block segmentation pattern defines a segmentation type; and a step of decoding the plurality of blocks; the set of block segmentation patterns includes a first block segmentation pattern and a second block segmentation pattern that respectively define a segmentation direction and a segmentation number; in the step of performing the above segmentation, the use of the second block segmentation pattern with a segmentation number of 2 is restricted.
[0445] The decoding method according to an embodiment of the present disclosure may also include: a step of using a set of block segmentation patterns obtained by combining one or more block segmentation patterns to segment an encoded signal read from a memory into a set of blocks composed of a plurality of blocks, where the block segmentation pattern defines a segmentation type; and a step of decoding the plurality of blocks; in the step of performing the segmentation, when the first set of blocks obtained by using the first set of block segmentation patterns is the same as the second set of blocks obtained by using the second set of block segmentation patterns, only one of the first set of block segmentation patterns or the second set of block segmentation patterns is used for segmentation.
[0446] The picture compression program according to an embodiment of the present disclosure may also include: using a set of block segmentation patterns obtained by combining one or more block segmentation patterns to segment a picture read from a memory into a plurality of blocks, where the block segmentation pattern defines a segmentation type; decoding the plurality of blocks; the set of block segmentation patterns is composed of a first block segmentation pattern and a second block segmentation pattern, the first block segmentation pattern defines a segmentation direction and a segmentation number for segmenting a first block, and the second block segmentation pattern defines a segmentation direction and a segmentation number for segmenting a second block, which is one of the blocks obtained after the segmentation of the first block; in the above segmentation, when the segmentation number of the first block segmentation pattern is 3, the second block is the central block among the blocks obtained after the segmentation of the first block, and the segmentation direction of the second block segmentation pattern is the same as the segmentation direction of the first block segmentation pattern, the second block segmentation pattern only includes a block segmentation pattern with a segmentation number of 3.
[0447] The picture compression program according to an embodiment of the present disclosure may also include: a step of using a set of block segmentation patterns obtained by combining one or more block segmentation patterns to segment a picture read from a memory into a plurality of blocks, where the block segmentation pattern defines a segmentation type; and a step of encoding the plurality of blocks; the set of block segmentation patterns is composed of a first block segmentation pattern and a second block segmentation pattern, the first block segmentation pattern defines a segmentation direction and a segmentation number for segmenting a first block, and the second block segmentation pattern defines a segmentation direction and a segmentation number for segmenting a second block, which is one of the blocks obtained after the segmentation of the first block; in the step of performing the segmentation, when the segmentation number of the first block segmentation pattern is 3, the second block is the central block among the blocks obtained after the segmentation of the first block, and the segmentation direction of the second block segmentation pattern is the same as the segmentation direction of the first block segmentation pattern, the second block segmentation pattern with a segmentation number of 2 is not used.
[0448] The picture compression program according to an embodiment of the present disclosure may also have: a step of dividing a picture read from a memory into a plurality of blocks using a set of block division patterns obtained by combining one or more block division patterns, where the block division patterns define division types; and a step of encoding the plurality of blocks; the set of block division patterns includes a first block division pattern and a second block division pattern that respectively define a division direction and a division quantity; in the step of performing the division, use of the second block division pattern with the division quantity of 2 is restricted.
[0449] The picture compression program according to an embodiment of the present disclosure may also have: a step of dividing a picture read from a memory into a set of blocks composed of a plurality of blocks using a set of block division patterns obtained by combining one or more block division patterns, where the block division patterns define division types; and a step of encoding the plurality of blocks; in the step of performing the division, when the first block set obtained using the first set of block division patterns is the same as the second block set obtained using the second set of block division patterns, only one of the first set of block division patterns or the second set of block division patterns is used for division.
[0450] Industrial applicability
[0451] It can be used for encoding / decoding of multimedia data, particularly for image and video encoding / decoding devices using block encoding / decoding.
[0452] Reference numeral description
[0453] 100 Encoding device
[0454] 102 Division unit
[0455] 104 Subtraction unit
[0456] 106, 5001 Transformation unit
[0457] 108, 5002 Quantization unit
[0458] 110, 5009 Entropy encoding unit
[0459] 112, 5003, 6002 Inverse quantization unit
[0460] 114, 5004, 6003 Inverse transformation unit
[0461] 116 Addition unit
[0462] 118, 5005, 6004 Block memory
[0463] 120 Loop filtering unit
[0464] 122, 5006, 6005 Frame memory
[0465] 124, 5007, 6006 Intra prediction unit
[0466] 126, 5008, 6007 Inter prediction unit
[0467] 128 Prediction control unit
[0468] 200 Decoding device
[0469] 202, 6001 Entropy decoding unit
[0470] 204 Inverse quantization unit
[0471] 206 Inverse transform unit
[0472] 208 Addition unit
[0473] 210 Block memory
[0474] 212 Loop filtering unit
[0475] 214 Frame memory
[0476] 216 Intra prediction unit
[0477] 218 Inter prediction unit
[0478] 220 Prediction control unit
[0479] 5000 Video encoding device
[0480] 5010, 6008 Block segmentation decision unit
[0481] 6000 Video decoding device
Claims
1. An encoding device that encodes pictures, wherein, Comprising: A processor; and A memory; The above-mentioned processor has: A block splitting determination unit that uses a set of block splitting patterns obtained by combining one or more block splitting patterns to split the above-mentioned picture read from the above-mentioned memory into a plurality of blocks, and the above-mentioned block splitting pattern defines a splitting type; and An encoding unit that encodes the above-mentioned plurality of blocks; The above-mentioned set of block splitting patterns consists of a first block splitting pattern and a second block splitting pattern. The above-mentioned first block splitting pattern defines a splitting direction and a splitting number for splitting a first block, and the above-mentioned second block splitting pattern defines a splitting direction and a splitting number for splitting a second block, which is one of the blocks obtained after splitting the above-mentioned first block; The parameter for identifying the above-mentioned second block splitting pattern includes a first flag indicating in which of the horizontal direction and the vertical direction the above-mentioned second block is split; In the above-mentioned block splitting determination unit, When the above-mentioned splitting number of the above-mentioned first block splitting pattern is 3, the above-mentioned second block is the central block among the blocks obtained after splitting the above-mentioned first block, and the above-mentioned splitting direction of the above-mentioned second block splitting pattern indicated by the above-mentioned first flag is the same as the above-mentioned splitting direction of the above-mentioned first block splitting pattern, the above-mentioned second block splitting pattern only includes a block splitting pattern with a splitting number of 3; When the above-mentioned splitting number of the above-mentioned first block splitting pattern is 3, the above-mentioned second block is the central block among the blocks obtained after splitting the above-mentioned first block, and the above-mentioned splitting direction of the above-mentioned second block splitting pattern indicated by the above-mentioned first flag is different from the above-mentioned splitting direction of the above-mentioned first block splitting pattern, the above-mentioned second block splitting pattern includes a block splitting pattern with a splitting number of 2.
2. A decoding device that decodes an encoded signal, wherein, Comprising: A processor; and A memory; The above-mentioned processor has: A block splitting determination unit that uses a set of block splitting patterns obtained by combining one or more block splitting patterns to split the above-mentioned encoded signal read from the above-mentioned memory into a plurality of blocks, and the above-mentioned block splitting pattern defines a splitting type; and A decoding unit that decodes the above-mentioned plurality of blocks; The above-mentioned set of block splitting patterns consists of a first block splitting pattern and a second block splitting pattern. The above-mentioned first block splitting pattern defines a splitting direction and a splitting number for splitting a first block, and the above-mentioned second block splitting pattern defines a splitting direction and a splitting number for splitting a second block, which is one of the blocks obtained after splitting the above-mentioned first block; The parameter for identifying the above-mentioned second block splitting pattern includes a first flag indicating in which of the horizontal direction and the vertical direction the above-mentioned second block is split; In the above-mentioned block splitting determination unit, When the above-mentioned splitting number of the above-mentioned first block splitting pattern is 3, the above-mentioned second block is the central block among the blocks obtained after splitting the above-mentioned first block, and the above-mentioned splitting direction of the above-mentioned second block splitting pattern indicated by the above-mentioned first flag is the same as the above-mentioned splitting direction of the above-mentioned first block splitting pattern, the above-mentioned second block splitting pattern only includes a block splitting pattern with a splitting number of 3; When the number of divisions in the above-described first block division pattern is 3, the second block is the central block among the blocks obtained after the division of the first block, and the division direction of the second block division pattern indicated by the first flag is different from the division direction of the first block division pattern, the second block division pattern includes a block division pattern with a division number of 2.
Citation Information
Patent Citations
Method and apparatus for encoding / decoding video signals
CN107439014A
Color image encoding device, color image decoding device, color image encoding method and color image decoding method
JP2014204311A