Image encoding device, image encoding device, image encoding method, and image encoding method

By determining deblocking filter intensity based on prediction modes for intra-inter-prediction, the encoding process is optimized for mixed intra-inter-prediction, improving encoding efficiency and image quality.

JP7867352B2Active Publication Date: 2026-05-29CANON KK

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
CANON KK
Filing Date
2022-03-22
Publication Date
2026-05-29

Smart Images

  • Figure 0007867352000001
    Figure 0007867352000001
  • Figure 0007867352000002
    Figure 0007867352000002
  • Figure 0007867352000003
    Figure 0007867352000003
Patent Text Reader

Abstract

To provide a technique to allow deblocking filter processing designed for intra-inter mixed prediction.SOLUTION: When at least one of a first block and a second block is a block applying a prediction mode for deriving a predicted pixel of a block to be encoded by using pixels in an image including the block, an image encoding device determines the intensity of deblocking filter processing performed on the boundary between the first block and the second block as a first intensity. When at least one of the first block and the second block is a block applying a prediction mode for deriving a predicted pixel by using pixels in an image including a block to be encoded for one area in the block, and deriving a predicted pixel by using pixels in the other image different from the image including the block for the other area different from one area in the block, the image encoding device determines the intensity as an intensity based on the first intensity.SELECTED DRAWING: Figure 13
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to image encoding technology and image decoding technology.

Background Art

[0002] As an encoding method for compression recording of moving images, the VVC (Versatile Video Coding) encoding method (hereinafter referred to as VVC) is known. In VVC, in order to improve the encoding efficiency, a basic block of up to 128×128 pixels is divided into sub-blocks not only in the conventional square shape but also in a rectangular shape.

[0003] Also, in VVC, a matrix called a quantization matrix is used to weight the coefficients after orthogonal transformation (hereinafter referred to as orthogonal transformation coefficients) according to frequency components. By further reducing the data of high-frequency components that are less noticeable in human vision degradation, it is possible to improve the compression efficiency while maintaining the image quality. Patent Document 1 discloses a technique for encoding such a quantization matrix.

[0004] Also, in VVC, by performing adaptive deblocking filter processing on the block boundary of the reconstructed image obtained by adding the signal after inverse quantization and inverse transformation processing and the predicted image, it is possible to suppress block distortion that is easily noticeable to human vision and prevent the propagation of image quality degradation to the predicted image. Patent Document 2 discloses a technique related to such a deblocking filter.

[0005] In recent years, the JVET (Joint Video Experts Team) that has standardized VVC has been conducting technical studies to achieve a compression efficiency higher than that of VVC. In order to improve the encoding efficiency, in addition to the conventional intra prediction and inter prediction, a new prediction method in which intra prediction pixels and inter prediction pixels are mixed within the same sub-block (hereinafter referred to as intra-inter mixed prediction) is being studied.

Prior Art Documents

Patent Documents

[0006] [Patent Document 1] Japanese Patent Publication No. 2013-38758 [Patent Document 2] Special Publication No. 2014-507863 [Overview of the project] [Problems that the invention aims to solve]

[0007] Deblocking filters in VVC are based on conventional prediction methods such as intra-prediction and inter-prediction, and are not compatible with the newer prediction method of mixed intra-inter-prediction. Therefore, this invention provides a technology to enable deblocking filter processing that is compatible with mixed intra-inter-prediction. [Means for solving the problem]

[0008] One aspect of the present invention is prediction in block units. Using the mode, images Encoding means for encoding, The first block in the aforementioned image, The aforementioned In images The aforementioned The boundary between the first block and the adjacent second block Executed The intensity of the deblocking filter process decision do decision means and The aforementioned decision means Determined by Strength Therefore, the boundary Deblocking filter process Execute Processing means and Equipped with, The encoding means predicts the target block mode as, The aforementioned Using pixels in the image containing the target block, The aforementioned Predicted pixels of the target block but Derivation It will be done First prediction mode, The aforementioned Using pixels from another image different from the image containing the target block, The aforementioned Predicted pixels of the target block but Derivation It will be done The second prediction mode, and The aforementioned For the region in the target block 1 Predicted pixels are derived using the pixels in the image including the target block Without using pixels in another image different from the image containing the target block, For the region in the target block but Derivation So , The aforementioned For the region in the target block that is 1 Different from the above 2nd Predicted pixels are derived using the pixels in another image different from the image including the target block Without using pixels in the image including the target block For the region in the target block that is but Derivation It will be done The third prediction mode, and One of a plurality of prediction modes including can be used, Each of the first and second regions in the target block can have a shape different from a rectangle. The decision means is When at least one of the first block and the second block is a block to which the first prediction mode is applied but Apply done For the boundary, the first intensity is used as the intensity related to the deblocking filter processing and then executed For the boundary, the first intensity is used as the intensity related to the deblocking filter processing decision And When at least one of the first block and the second block is a block to which the third prediction mode is applied but Apply done For the boundary, the second intensity is used as the intensity related to the deblocking filter processing and then executed For the boundary, the second intensity is used as the intensity related to the deblocking filter processing decision Do It is characterized by the above.

Effect of the Invention

[0009] According to the present invention, it is possible to provide a technique for enabling deblocking filter processing corresponding to intra / inter mixed prediction.

Brief Description of the Drawings

[0010] [Figure 1] Block diagram showing a functional configuration example of an image encoding device. [Figure 2] A block diagram showing an example of the functional configuration of an image decoding device. [Figure 3] A flowchart of the encoding process performed by an image encoding device. [Figure 4] A flowchart of the decoding process in an image decoding device. [Figure 5] A block diagram showing an example of a computer device hardware configuration. [Figure 6] A diagram showing an example of a bitstream configuration. [Figure 7] A diagram illustrating an example of how to divide a basic block of 700 into subblocks. [Figure 8] A diagram showing an example of a quantization matrix. [Figure 9] A diagram showing the reference order of the values ​​of each element in a quantization matrix. [Figure 10] A figure showing an example of a one-dimensional array. [Figure 11] A diagram showing an example of an encoding table. [Figure 12] A diagram illustrating the prediction of a mixed intranet and internet environment. [Figure 13] A block diagram showing an example of the functional configuration of an image encoding device. [Figure 14] A diagram illustrating the pixel group at the boundary between subblock P and subblock Q. [Figure 15] A flowchart of the encoding process performed by an image encoding device. [Figure 16] A block diagram showing an example of the functional configuration of an image decoding device. [Figure 17] A flowchart of the decoding process in an image decoding device. [Modes for carrying out the invention]

[0011] The embodiments will be described in detail below with reference to the attached drawings. Note that the following embodiments do not limit the invention to the claims. While the embodiments describe multiple features, not all of these features are essential to the invention, and the features may be combined in any way. Furthermore, in the attached drawings, the same or similar configurations are given the same reference numerals, and redundant descriptions are omitted.

[0012] [First Embodiment] The image encoding device according to this embodiment applies an intra-predicted image obtained by intra-prediction to a portion of the block to be encoded contained in the image, and applies an inter-predicted image obtained by inter-prediction to other portions of the block that differ from the portion in question, thereby obtaining a predicted image. The image encoding device then encodes the quantization coefficients obtained by quantizing the orthogonal transformation coefficients of the difference between the block and the predicted image using a quantization matrix (first encoding).

[0013] First, an example of the functional configuration of the image encoding device according to this embodiment will be explained using the block diagram in Figure 1. The control unit 150 controls the operation of the entire image encoding device. The splitting unit 102 divides the input image into a plurality of basic blocks and outputs each of the divided basic blocks. The input image may be an image of each frame that makes up a moving image (for example, an image of each frame in a moving image with 30 frames / second), or it may be a still image that is captured periodically or irregularly. Furthermore, the splitting unit 102 may acquire the input image from any device; for example, it may acquire it from an imaging device such as a video camera, from a device that holds multiple images, or from memory that the device can access.

[0014] The data storage unit 103 holds quantization matrices corresponding to each of the multiple prediction processes. In this embodiment, the data storage unit 103 holds a quantization matrix corresponding to intra-prediction, which is an in-frame prediction; a quantization matrix corresponding to inter-prediction, which is an inter-frame prediction; and a quantization matrix corresponding to the above-mentioned intra-inter-mixed prediction. Each quantization matrix held by the data storage unit 103 may be a quantization matrix with default element values, or it may be a quantization matrix generated by the control unit 150 in response to user operation. Furthermore, each quantization matrix held by the data storage unit 103 may be a quantization matrix generated by the control unit 150 according to the characteristics of the input image (such as the amount and frequency of edges contained in the input image).

[0015] The prediction unit 104 divides each basic block into multiple subblocks. The prediction unit 104 then obtains a predicted image for each subblock using either intra-prediction, inter-prediction, or a combination of intra- and inter-prediction, and calculates the difference between the subblock and the predicted image as the prediction error. The prediction unit 104 also generates prediction information, including information indicating the method of dividing the basic block, a prediction mode indicating the prediction for obtaining the predicted image of the subblock, and motion vectors, which are necessary for prediction.

[0016] The transformation / quantization unit 105 generates orthogonal transformation coefficients for each subblock by performing an orthogonal transformation (frequency transformation) on the prediction error of each subblock obtained by the prediction unit 104, obtains a quantization matrix from the holding unit 103 corresponding to the prediction (intra prediction, inter prediction, intra / inter mixed prediction) made by the prediction unit 104 to obtain the predicted image of the subblock, and generates quantization coefficients for the subblock (quantization result of the orthogonal transformation coefficients) by quantizing the orthogonal transformation coefficients using the obtained quantization matrix.

[0017] The inverse quantization / inverse transformation unit 106 generates orthogonal transformation coefficients by performing inverse quantization of the quantization coefficients of each subblock generated by the transformation / quantization unit 105 using the quantization matrix used by the transformation / quantization unit 105 to generate the quantization coefficients, and then generates (reconstructs) the prediction error by performing an inverse orthogonal transformation on these orthogonal transformation coefficients.

[0018] The image playback unit 107 generates a predicted image from the images stored in the frame memory 108 based on the prediction information generated by the prediction unit 104, and then reproduces the image from the predicted image and the prediction error generated by the inverse quantization / inverse transform unit 106. The image playback unit 107 then stores the reproduced image in the frame memory 108. The images stored in the frame memory 108 are the images that the prediction unit 104 refers to when making predictions about the image of the current frame or the next frame.

[0019] The in-loop filter unit 109 performs in-loop filtering, such as deblocking filtering and sample adaptive offsetting, on the images stored in the frame memory 108.

[0020] The encoding unit 110 encodes the quantization coefficients generated by the conversion / quantization unit 105 and the prediction information generated by the prediction unit 104 to generate encoded data (coded data).

[0021] The encoding unit 113 encodes the quantization matrix (including at least the quantization matrix used by the transformation / quantization unit 105 for quantization) held in the holding unit 103 to generate encoded data (coded data).

[0022] The integrated encoding unit 111 generates header code data using the encoded data generated by the encoding unit 113, and outputs a bitstream that includes the encoded data generated by the encoding unit 110 and the header code data.

[0023] Furthermore, the output destination of the bitstream is not limited to a specific destination. For example, the bitstream may be output to the memory of the image encoding device, to an external device via the network to which the image encoding device is connected, or transmitted externally for broadcasting purposes.

[0024] Next, the operation of the image encoding device according to this embodiment will be described. First, the encoding of the input image will be described. The division unit 102 divides the input image into a plurality of basic blocks and outputs each of the divided basic blocks.

[0025] The prediction unit 104 divides each basic block into multiple subblocks. Figures 7(a) to 7(f) show an example of how to divide a basic block 700 into subblocks.

[0026] Figure 7(a) shows the basic 8x8 pixel block 700 (= subblock) which has not been divided into subblocks. Figure 7(b) shows an example of a conventional square subblock division, in which the basic 8x8 pixel block 700 is divided into four 4x4 pixel subblocks (quadtree division).

[0027] Figures 7(c) to 7(f) show examples of rectangular subblock partitioning. In Figure 7(c), the 8x8 basic block 700 is divided into two 4-pixel (horizontal) x 8-pixel (vertical) subblocks (binary tree partitioning). Similarly, in Figure 7(d), the 8x8 basic block 700 is divided into two 8-pixel (horizontal) x 4-pixel (vertical) subblocks (binary tree partitioning).

[0028] In Figure 7(e), the 8x8 basic block 700 is divided into three subblocks: a 2x8 pixel subblock (horizontal) x 2x8 pixel subblock (vertical), a 4x8 pixel subblock (horizontal) x 2x8 pixel subblock (vertical). In other words, in Figure 7(e), the basic block 700 is divided into subblocks with a width (horizontal length) in a ratio of 1:2:1 (ternary tree division).

[0029] In Figure 7(f), the basic 8x8 block 700 is divided into three subblocks: an 8-pixel (horizontal) x 2-pixel (vertical) subblock, an 8-pixel (horizontal) x 4-pixel (vertical) subblock, and an 8-pixel (horizontal) x 2-pixel (vertical) subblock. In other words, in Figure 7(f), the basic block 700 is divided into subblocks with a height (vertical length) in a ratio of 1:2:1 (ternary tree division).

[0030] Thus, in this embodiment, encoding is performed using not only square subblocks but also rectangular subblocks. In this embodiment, prediction information is generated that includes information indicating how to divide the basic blocks in this way. Note that the division method shown in Figure 7 is merely one example, and the method of dividing the basic blocks into subblocks is not limited to the division method shown in Figure 7.

[0031] The prediction unit 104 then determines the prediction (prediction mode) to be performed for each subblock. For each subblock, the prediction unit 104 generates a predicted image based on the prediction mode determined for that subblock and the encoded pixels, and calculates the difference between the subblock and the predicted image as the prediction error. The prediction unit 104 also generates prediction information, which includes information indicating the division method of the basic block, the prediction mode of the subblock, and motion vectors, as "information necessary for prediction".

[0032] Here, the predictions used in this embodiment will be explained again. In this embodiment, three types of predictions (prediction modes) are used: intra prediction, inter prediction, and intra / inter mixed prediction.

[0033] In intra-prediction (first prediction mode), predicted pixels for a block to be encoded (a subblock in this embodiment) are generated using encoded pixels located spatially around the block. In other words, intra-prediction generates predicted pixels (predicted images) for a block to be encoded using encoded pixels within the frame (image) containing the block to be encoded. For subblocks on which intra-prediction has been performed, information indicating the intra-prediction method, such as horizontal prediction, vertical prediction, or DC prediction, is generated as "information necessary for prediction."

[0034] In interpretation (second prediction mode), the predicted pixels of the target block (subblock in this embodiment) are generated using encoded pixels from other frames (other images) that are (temporarily) different from the frame (image) to which the target block belongs. For subblocks that have undergone interpretation, motion information indicating the reference frame and motion vector is generated as "information necessary for prediction".

[0035] In intra / inter mixed prediction (third prediction mode), the target block to be encoded (a subblock in this embodiment) is first divided into two divided regions by a diagonal line segment. Then, the predicted pixels of one divided region are obtained as "predicted pixels obtained for the one divided region by intra prediction for the target block to be encoded." Similarly, the predicted pixels of the other divided region are obtained as "predicted pixels obtained for the other divided region by inter prediction for the target block to be encoded." In other words, the predicted pixels of one divided region in the predicted image obtained by intra / inter mixed prediction for the target block to be encoded are "predicted pixels obtained for the one divided region by intra prediction for the target block to be encoded." Similarly, the predicted pixels of the other divided region in the predicted image obtained by intra / inter mixed prediction for the target block to be encoded are "predicted pixels obtained for the other divided region by inter prediction for the target block to be encoded." An example of the division of the target block to be encoded in intra / inter mixed prediction is shown in Figure 12.

[0036] As shown in Figure 12(a), suppose the encoding target block 1200 is divided into divided region 1200a and divided region 1200b by a line segment passing through the vertex of the upper left corner and the vertex of the lower right corner of the encoding target block 1200. The intra / inter mixed prediction process for the encoding target block 1200 by the prediction unit 104 in this case will be explained with reference to Figures 12(a) to (d). At this time, the prediction unit 104 generates an intra-predicted image 1201 (Figure 12(b)) by performing intra-prediction on the encoding target block 1200. Here, the intra-predicted image 1201 includes region 1201a, which is in the same position as divided region 1200a, and region 1201b, which is in the same position as divided region 1200b. The prediction unit 104 then defines the intra-predicted pixels (intra-predicted pixels) included in the intra-predicted image 1201 that belong to region 1201a, which is in the same position as the divided region 1200a, as "predicted pixels of divided region 1200a". It also generates an inter-predicted image 1202 (Figure 12(c)) by performing inter-prediction on the encoding target block 1200. Here, the inter-predicted image 1202 includes region 1202a, which is in the same position as the divided region 1200a, and region 1202b, which is in the same position as the divided region 1200b. The prediction unit 104 then defines the inter-predicted pixels (intra-predicted pixels) included in the inter-predicted image 1202 that belong to region 1202b, which is in the same position as the divided region 1200b, as "predicted pixels of divided region 1200b". The prediction unit 104 then generates a predicted image 1203 (Figure 12(d)) consisting of intra-predicted pixels included in region 1201a, which is the "predicted pixels of divided region 1200a", and inter-predicted pixels included in region 1202b, which is the "predicted pixels of divided region 1200b".

[0037] As described above, the prediction unit 104 generates an intra-predicted image 1201 (Figure 12(b)) by intra-prediction on the target block 1200, and further generates an inter-predicted image 1202 (Figure 12(c)) by inter-prediction on the target block 1200. The prediction unit 104 then places the intra-predicted pixels in the intra-predicted image 1201 at coordinates (x, y) included in region 1201a corresponding to the divided region 1200a at the same coordinates (x, y) in the predicted image 1203. The prediction unit 104 also places the inter-predicted pixels in the inter-predicted image 1202 at coordinates (x, y) included in region 1202b corresponding to the divided region 1200b at the same coordinates (x, y) in the predicted image 1203. In this way, the predicted image 1203 shown in Figure 12(d) is generated.

[0038] Next, with reference to Figures 12(e) to (h), the intra-inter-mixed prediction process performed by the prediction unit 104 for the encoding target block 1200 will be explained. In this example, as shown in Figure 12(e), the encoding target block 1200 is divided into divided region 1200c and divided region 1200d by a line segment passing through the midpoint between the vertex of the upper left corner and the vertex of the lower left corner and the vertex of the upper right corner. At this time, the prediction unit 104 generates an intra-predicted image 1201 (Figure 12(f)) by performing intra-prediction on the encoding target block 1200. Here, the intra-predicted image 1201 includes region 1201c, which is in the same position as divided region 1200c, and region 1201d, which is in the same position as divided region 1200d. The prediction unit 104 then defines the intra-predicted pixels (intra-predicted pixels) included in the intra-predicted image 1201 that belong to region 1201c, which is in the same position as the divided region 1200c, as "predicted pixels of divided region 1200c". The prediction unit 104 also generates an inter-predicted image 1202 (Figure 12(g)) by performing inter-prediction on the encoding target block 1200. Here, the inter-predicted image 1202 includes region 1202c, which is in the same position as the divided region 1200c, and region 1202d, which is in the same position as the divided region 1200d. The prediction unit 104 then defines the inter-predicted pixels (intra-predicted pixels) included in the inter-predicted image 1202 that belong to region 1202d, which is in the same position as the divided region 1200d, as "predicted pixels of divided region 1200d". The prediction unit 104 then generates a predicted image 1203 (Figure 12(h)) consisting of intra-predicted pixels included in region 1201c, which is the "predicted pixels of divided region 1200c," and inter-predicted pixels included in region 1202d, which is the "predicted pixels of divided region 1200d."

[0039] As described above, the prediction unit 104 generates an intra-predicted image 1201 (Figure 12(f)) by intra-prediction on the target block 1200, and further generates an inter-predicted image 1202 (Figure 12(g)) by inter-prediction on the target block 1200. The prediction unit 104 then places the intra-predicted pixels in the intra-predicted image 1201 that are located at coordinates (x, y) included in region 1201c corresponding to the divided region 1200c at the same coordinates (x, y) in the predicted image 1203. The prediction unit 104 also places the inter-predicted pixels in the inter-predicted image 1202 that are located at coordinates (x, y) included in region 1202d corresponding to the divided region 1200d at the same coordinates (x, y) in the predicted image 1203. By doing so, the predicted image 1203 shown in Figure 12(d) is generated.

[0040] For subblocks that perform intra-internal mixed prediction, information is generated as "information necessary for prediction," including information indicating the intra-prediction method, motion information indicating the reference frames and motion vectors, and information defining the division region (for example, the information defining the line segment mentioned above).

[0041] The prediction unit 104 determines the prediction mode of the subblock of interest by the following process: The prediction unit 104 generates a difference image between the predicted image generated by intra-prediction for the subblock of interest and the subblock of interest. The prediction unit 104 also generates a difference image between the predicted image generated by inter-prediction for the subblock of interest and the subblock of interest. The prediction unit 104 also generates a difference image between the predicted image generated by a mixed intra-inter-prediction for the subblock of interest and the subblock of interest. The pixel value at pixel position (x, y) in the difference image C of image A and image B is the difference between the pixel value AA at pixel position (x, y) in image A and the pixel value BB at pixel position (x, y) in image B (such as the absolute value of the difference between AA and BB, or the squared value of the difference between AA and BB). The prediction unit 104 then identifies the prediction image in which the sum of the pixel values ​​of all pixels in the difference image is smallest, and determines the prediction made to the subblock of interest in order to obtain that prediction image as the "prediction mode of the subblock of interest". Note that the method for determining the prediction mode of the subblock of interest is not limited to the method described above.

[0042] The prediction unit 104 then determines the prediction image generated in the prediction mode determined for each subblock as the "prediction image of the subblock" and generates a prediction error from the subblock and the prediction image. The prediction unit 104 also generates prediction information for each subblock, including the prediction mode determined for the subblock and the "information necessary for prediction" generated for the subblock.

[0043] The transformation / quantization unit 105 applies an orthogonal transformation process to the prediction error of each subblock, corresponding to the size of the prediction error, to generate orthogonal transformation coefficients. Then, for each subblock, the transformation / quantization unit 105 obtains the quantization matrix corresponding to the prediction mode of the subblock from the quantization matrix held by the holding unit 103, and uses the obtained quantization matrix to quantize the orthogonal transformation coefficients of the subblock to generate quantization coefficients.

[0044] For example, suppose the storage unit 103 holds an 8x8 quantization matrix (all 64 elements are quantization step values) as exemplified in Figure 8(a) for use in quantizing the orthogonal transformation coefficients of the prediction error obtained when intra prediction is performed on an 8x8 subblock. Also, suppose the storage unit 103 holds an 8x8 quantization matrix (all 64 elements are quantization step values) as exemplified in Figure 8(b) for use in quantizing the orthogonal transformation coefficients of the prediction error obtained when inter prediction is performed on an 8x8 subblock. Also, suppose the storage unit 103 holds an 8x8 quantization matrix (all 64 elements are quantization step values) as exemplified in Figure 8(c) for use in quantizing the orthogonal transformation coefficients of the prediction error obtained when intra / inter mixed prediction is performed on an 8x8 subblock.

[0045] In this case, the transformation and quantization unit 105 quantizes the orthogonal transformation coefficients of the "prediction error obtained by intra-prediction for an 8x8 pixel subblock" using the intra-prediction quantization matrix shown in Figure 8(a).

[0046] Furthermore, the conversion and quantization unit 105 quantizes the orthogonal transformation coefficients of the "prediction error obtained by interprediction for an 8x8 pixel subblock" using the interprediction quantization matrix shown in Figure 8(b).

[0047] Furthermore, the conversion and quantization unit 105 quantizes the orthogonal transformation coefficients of the "prediction error obtained by intra-inter mixed prediction for an 8x8 pixel subblock" using the quantization matrix for intra-inter mixed prediction shown in Figure 8(c).

[0048] The inverse quantization / inverse transformation unit 106 generates orthogonal transformation coefficients by performing inverse quantization on the quantization coefficients of each subblock generated by the transformation / quantization unit 105 using the quantization matrix used by the transformation / quantization unit 105 for the quantization of the subblock, and then generates (reconstructs) prediction errors by performing an inverse orthogonal transformation on these orthogonal transformation coefficients.

[0049] The image playback unit 107 generates a predicted image from the images stored in the frame memory 108 based on the prediction information generated by the prediction unit 104, and then adds the predicted image to the prediction error generated (played back) by the inverse quantization / inverse transform unit 106 to play back the image of the subblock. The image playback unit 107 then stores the played back image in the frame memory 108.

[0050] The in-loop filter unit 109 performs in-loop filtering, such as deblocking filtering and sample adaptive offsetting, on the image stored in the frame memory 108, and stores the in-loop filtered image in the frame memory 108.

[0051] The encoding unit 110 generates encoded data by entropy encoding the quantization coefficients of the subblock generated by the transformation / quantization unit 105 and the prediction information of the subblock generated by the prediction unit 104 for each subblock. While no specific method is designated for entropy encoding, Golomb coding, arithmetic coding, Huffman coding, etc., can be used.

[0052] Next, the encoding of the quantization matrix will be explained. The quantization matrix held by the storage unit 103 is generated according to the size of the subblock to be encoded and the prediction mode. For example, as shown in Figure 7, when adopting sizes such as 8 pixels x 8 pixels, 4 pixels x 4 pixels, 8 pixels x 4 pixels, 4 pixels x 8 pixels, 8 pixels x 2 pixels, and 2 pixels x 8 pixels as the size of the subblock to be divided, the quantization matrix of the adopted size is registered in the storage unit 103. In addition, quantization matrices are prepared for intra-prediction, inter-prediction, and intra-inter-mixed prediction, and are registered in the storage unit 103.

[0053] The method for generating a quantization matrix according to the subblock size and prediction mode is not limited to a specific method as described above, nor is the method for managing such a quantization matrix in the holding unit 103 limited to a specific method.

[0054] In this embodiment, the quantization matrix held by the holding unit 103 is assumed to be held in a two-dimensional shape as shown in Figure 8, but the elements within the quantization matrix are not limited to this. Furthermore, it is possible to hold multiple quantization matrices for the same prediction method depending on the size of the subblock or whether the encoding target is a luminance block or a chrominance block. Generally, in order to realize quantization processing according to the characteristics of human vision, the elements for the DC component corresponding to the upper left corner of the quantization matrix are small, and the elements for the AC component corresponding to the lower right corner are large, as shown in Figure 8.

[0055] The encoding unit 113 reads the quantization matrix (including at least the quantization matrix used by the transformation / quantization unit 105 for quantization) held in the holding unit 103 and encodes the read quantization matrix. For example, the encoding unit 113 encodes the quantization matrix of interest using the following process.

[0056] The encoding unit 113 references the values ​​of each element in the two-dimensional array, the quantization matrix of interest, in a predetermined order and generates a one-dimensional array by listing the difference between the value of the currently referenced element and the value of the previously referenced element. For example, if the quantization matrix in Figure 8(c) is the quantization matrix of interest, the encoding unit 113 references the values ​​of each element from the element in the upper left corner to the element in the lower right corner of the quantization matrix of interest in the order indicated by the arrows, as shown in Figure 9.

[0057] In this case, the value of the first element referenced is "8", and the value of the element referenced immediately before does not exist, so the encoding unit 113 outputs a predetermined value or a value obtained by some method as the output value. For example, the encoding unit 113 may output the value of the currently referenced element, "8", as the output value, or it may output a value obtained by subtracting a predetermined value from the value of the element, "8", and the output value does not have to be a value determined by a specific method.

[0058] The next element to be referenced has a value of "11," and the previously referenced element has a value of "8." Therefore, the encoding unit 113 outputs the difference value "+3" obtained by subtracting the value of the previously referenced element "8" from the value of the currently referenced element "11" as the output value. In this way, the encoding unit 113 references the values ​​of each element in the quantization matrix in a predetermined order, obtains and outputs the output values, and generates a one-dimensional array by arranging the output values ​​in the order of output.

[0059] Figure 10(a) shows the one-dimensional array generated from the quantization matrix in Figure 8(a) by this process. Figure 10(b) shows the one-dimensional array generated from the quantization matrix in Figure 8(b) by this process. Figure 10(c) shows the one-dimensional array generated from the quantization matrix in Figure 8(c) by this process. In Figure 10, it is assumed that the default value is set to "8".

[0060] The encoding unit 113 then encodes the one-dimensional array generated for the quantization matrix of interest. For example, the encoding unit 113 refers to the encoding table exemplified in Figure 11(a) and generates a bit sequence as encoded data by replacing each element value in the one-dimensional array with the corresponding binary code. Note that the encoding table is not limited to the encoding table shown in Figure 11(a); for example, the encoding table exemplified in Figure 11(b) may also be used.

[0061] Returning to Figure 1, the integrated encoding unit 111 integrates "header information necessary for image encoding" into the encoded data generated by the encoding unit 113, and generates header code data using the encoded data with the integrated header information. The integrated encoding unit 111 then generates and outputs a bitstream by multiplexing the encoded data generated by the encoding unit 110 and the header code data.

[0062] Figure 6(a) shows an example of the configuration of a bitstream generated by the integrated encoding unit 111. The sequence header contains the encoded data of the quantization matrix and the encoded data of each element. However, the location of encoding is not limited to this, and it is also possible to encode in the picture header or other headers. Furthermore, if the quantization matrix is ​​to be changed within a single sequence, it is possible to update it by re-encoding the quantization matrix. In this case, the entire quantization matrix may be rewritten, or it is possible to change only a part of it by specifying the prediction mode of the quantization matrix corresponding to the quantization matrix to be rewritten.

[0063] The encoding process performed by the image encoding device described above will now be explained according to the flowchart in Figure 3. Note that the process according to the flowchart in Figure 3 is for encoding a single input image. Therefore, when encoding images of each frame in a video, or multiple images captured periodically or irregularly, the process from steps S304 to S311 will be repeated for each image.

[0064] Furthermore, it is assumed that, prior to the start of the process according to the flowchart in Figure 3, the holding unit 103 already has registered the quantization matrix corresponding to intra prediction, the quantization matrix corresponding to inter prediction, and the quantization matrix corresponding to the intra / inter mixed prediction. As described above, all of the quantization matrices held by the holding unit 103 are quantization matrices corresponding to the size of the subblocks to be divided.

[0065] In step S302, the encoding unit 113 reads out the quantization matrix (including at least the quantization matrix used by the transformation / quantization unit 105 for quantization) held in the holding unit 103, encodes the read-out quantization matrix, and generates encoded data.

[0066] In step S303, the integrated encoding unit 111 generates "header information necessary for image encoding". The integrated encoding unit 111 then integrates the "header information necessary for image encoding" with the encoded data generated by the encoding unit 113 in step S302, and generates header code data using the encoded data with the integrated header information.

[0067] In step S304, the division unit 102 divides the input image into multiple basic blocks and outputs each of the divided basic blocks. Then, the prediction unit 104 divides each basic block into multiple subblocks.

[0068] In step S305, the prediction unit 104 selects one of the unselected subblocks in the input image as the selected subblock and determines the prediction mode for the selected subblock. The prediction unit 104 then performs a prediction on the selected subblock according to the determined prediction mode and obtains the predicted image, prediction error, and prediction information for the selected subblock.

[0069] In step S306, the transformation / quantization unit 105 applies an orthogonal transformation process corresponding to the size of the prediction error to the prediction error of the selected subblock obtained in step S305 to generate orthogonal transformation coefficients. The transformation / quantization unit 105 then obtains a quantization matrix from the quantization matrices held by the holding unit 103 that corresponds to the prediction mode of the selected subblock, and uses the obtained quantization matrix to quantize the orthogonal transformation coefficients of the subblock to obtain quantization coefficients.

[0070] In step S307, the inverse quantization / inverse transformation unit 106 generates orthogonal transformation coefficients by performing inverse quantization on the quantization coefficients of the selected subblock obtained in step S306 using the quantization matrix used by the transformation / quantization unit 105 for the selected subblock. Then, the inverse quantization / inverse transformation unit 106 generates (reconstructs) the prediction error by performing an inverse orthogonal transformation on the generated orthogonal transformation coefficients.

[0071] In step S308, the image playback unit 107 generates a predicted image from the images stored in the frame memory 108 based on the prediction information acquired in step S305, and plays back the image of the subblock by adding the predicted image to the prediction error generated in step S307. The image playback unit 107 then stores the played-back image in the frame memory 108.

[0072] In step S309, the coding unit 110 entropy encodes the quantization coefficients obtained in step S306 and the prediction information obtained in step S305 to generate coded data.

[0073] The integrated encoding unit 111 then generates and outputs a bitstream by multiplexing the header code data generated in step S303 and the encoded data generated by the encoding unit 110 in step S309.

[0074] In step S310, the control unit 150 determines whether all subblocks in the input image have been selected as selected subblocks. If, as a result of this determination, all subblocks in the input image have been selected as selected subblocks, the process proceeds to step S311. On the other hand, if there is one or more subblocks in the input image that have not yet been selected as selected subblocks, the process proceeds to step S305.

[0075] In step S311, the in-loop filter unit 109 performs in-loop filtering on the image stored in the frame memory 108 (the image of the selected subblock played back in step S308). The in-loop filter unit 109 then stores the in-loop filtered image in the frame memory 108.

[0076] This process allows the orthogonal transformation coefficients of the subblocks that underwent intra-interface mixed prediction to be quantized using the quantization matrix corresponding to the intra-interface mixed prediction. This enables control over quantization for each frequency component, thereby improving image quality.

[0077] <Variation> In the first embodiment, separate quantization matrices were prepared for intra-prediction, inter-prediction, and intra-inter-mixed prediction, and the quantization matrices corresponding to each prediction were encoded. However, some of these may be shared.

[0078] For example, when quantizing the orthogonal transformation coefficients of prediction errors obtained based on intra-internal mixed prediction, a quantization matrix corresponding to intra-prediction may be used instead of a quantization matrix corresponding to intra-internal mixed prediction. That is, for example, to quantize the orthogonal transformation coefficients of prediction errors obtained based on intra-internal mixed prediction, the quantization matrix for intra-prediction shown in Figure 8(a) may be used instead of the quantization matrix for intra-internal mixed prediction shown in Figure 8(c). In this case, encoding of the quantization matrix corresponding to intra-internal mixed prediction can be omitted. This reduces the amount of encoded data of the quantization matrix to be included in the bitstream, and also reduces image quality degradation caused by errors due to intra-prediction, such as block distortion.

[0079] Furthermore, when quantizing the orthogonal transformation coefficients of the prediction error obtained based on intra-internal mixed prediction, a quantization matrix corresponding to inter-prediction may be used instead of a quantization matrix corresponding to intra-internal mixed prediction. That is, for example, to quantize the orthogonal transformation coefficients of the prediction error obtained based on intra-internal mixed prediction, the quantization matrix for inter-prediction shown in Figure 8(b) may be used instead of the quantization matrix for intra-internal mixed prediction shown in Figure 8(c). In this case, the encoding of the quantization matrix corresponding to intra-internal mixed prediction can be omitted. This reduces the amount of encoded data of the quantization matrix to be included in the bitstream, and also reduces image quality degradation caused by errors due to inter-prediction, such as jerky motion.

[0080] Furthermore, in the predicted image of a subblock obtained by performing intra-inter prediction, the quantization matrix used for the subblock may be determined according to the respective sizes of the region of "predicted pixels obtained by intra prediction" and the region of "predicted pixels obtained by inter prediction".

[0081] For example, suppose that subblock 1200 is divided into divided region 1200c and divided region 1200d, as shown in Figure 12(e). Then, suppose that the predicted pixels for divided region 1200c are determined by "predicted pixels obtained by intra-prediction", and the predicted pixels for divided region 1200d are determined by "predicted pixels obtained by inter-prediction". Also, suppose that the size (area (number of pixels)) S1 of divided region 1200c and the size (area (number of pixels)) S2 of divided region 1200d are in a ratio of 1:3.

[0082] In this case, the size of the partitioned region 1200d to which interpretation is applied is larger than the size of the partitioned region 1200c to which intrapretation is applied in subblock 1200. Therefore, the transformation and quantization unit 105 applies a quantization matrix corresponding to interpretation (for example, the quantization matrix in Figure 8(b)) to the quantization of the orthogonal transformation coefficients of this subblock 1200.

[0083] Furthermore, if the size of the divided region 1200d to which interpretation is applied is smaller than the size of the divided region 1200c to which intrapretation is applied in subblock 1200, the transformation / quantization unit 105 applies a quantization matrix corresponding to intrapretation (for example, the quantization matrix in Figure 8(a)) to the quantization of the orthogonal transformation coefficients of subblock 1200. This reduces the image quality degradation of the larger divided region while omitting the encoding of the quantization matrix corresponding to the mixed intrapretation / interpretation. Therefore, the amount of encoded data of the quantization matrix to be included in the bitstream can be reduced.

[0084] Alternatively, a quantization matrix may be generated as the quantization matrix corresponding to intra-internal mixed prediction by combining the "quantization matrix corresponding to intra-prediction" and the "quantization matrix corresponding to inter-prediction" according to the ratio of S1 and S2. For example, the transformation / quantization unit 105 may generate the quantization matrix corresponding to intra-internal mixed prediction using the following equation (1).

[0085] QM[x][y]={w×QMinter[x][y]+(1-w)×QMintra[x][y]})…(1) Here, QM[x][y] represents the element value (quantization step value) at coordinate (x,y) in the quantization matrix corresponding to intra-inter mixed prediction. QMinter[x][y] represents the element value (quantization step value) at coordinate (x,y) in the quantization matrix corresponding to inter prediction. QMintra[x][y] represents the element value (quantization step value) at coordinate (x,y) in the quantization matrix corresponding to intra prediction. Furthermore, w is a value between 0 and 1 that indicates the proportion of the region in the subblock where inter prediction is used, and w = S2 / (S1+S2). In this way, the quantization matrix corresponding to intra-inter mixed prediction can be generated as needed and does not need to be created in advance, so the encoding of the quantization matrix can be omitted. Therefore, the amount of encoded data of the quantization matrix to be included in the bitstream can be reduced. In addition, appropriate quantization control can be performed according to the ratio of the sizes of the regions in which intra prediction and inter prediction are used, thereby improving image quality.

[0086] Furthermore, in the first embodiment, the quantization matrix applied to the subblock to which intra-inter mixed prediction is applied is configured to be uniquely determined, but it may also be configured to be selectable by introducing an identifier.

[0087] There are various methods for selecting the quantization matrix to be applied to a subblock to which intra-internal mixed prediction is applied, from a quantization matrix corresponding to intra-prediction, a quantization matrix corresponding to inter-prediction, and a quantization matrix corresponding to intra-internal mixed prediction. For example, the control unit 150 may select it according to user operation.

[0088] The bitstream then stores an identifier to identify the quantization matrix selected as the quantization matrix to be applied to the subblock to which intra-inter mixed prediction is applied.

[0089] For example, Figure 6(b) shows how the quantization matrix coding method information code is newly introduced as an identifier, allowing for the selective application of the quantization matrix to the subblock to which intra-inter mixed prediction is applied. For example, if the quantization matrix coding method information code is 0, it indicates that the quantization matrix corresponding to the intra prediction has been applied to the subblock to which intra-inter mixed prediction is applied. Also, if the quantization matrix coding method information code is 1, it indicates that the quantization matrix corresponding to the inter prediction has been applied to the subblock to which intra-inter mixed prediction is applied. On the other hand, if the quantization matrix coding method information code is 2, it indicates that the quantization matrix corresponding to the intra-inter mixed prediction has been applied to the subblock to which intra-inter mixed prediction is applied.

[0090] This makes it possible to selectively achieve a reduction in the amount of encoded data for the quantization matrix included in the bitstream, and to apply unique quantization control to subblocks using intra-inter-mixed prediction.

[0091] Furthermore, in the first embodiment, a prediction image was generated that included prediction pixels for one of the divided regions of the subblock (first prediction pixels) and prediction pixels for the other divided region (second prediction pixels). However, the method for generating the prediction image is not limited to this method. For example, in order to improve the image quality of the region near the boundary between one divided region and the other (boundary region), a third prediction pixel calculated by a weighted average of the first and second prediction pixels included in the boundary region may be used as the prediction pixel for the boundary region. In this case, the prediction pixel value of the corresponding region corresponding to one of the divided regions in the prediction image becomes the first prediction pixel, and the prediction pixel value of the corresponding region corresponding to the other divided region in the prediction image becomes the second prediction pixel. The prediction pixel value of the corresponding region corresponding to the boundary region in the prediction image becomes the third prediction pixel. This makes it possible to suppress the degradation of image quality in the boundary region of divided regions where different predictions are used, and to improve image quality.

[0092] Furthermore, in the first embodiment, three types of predictions were explained as examples: intra-prediction, inter-prediction, and intra-inter-mixed prediction. However, the types and number of predictions are not limited to these examples. For example, intra-inter-combined prediction (CIIP), which is used in VVC, may also be used. Intra-inter-combined prediction is a prediction method that calculates the pixels of the entire block to be encoded using a weighted average of predicted pixels by intra-prediction and predicted pixels by inter-prediction. In this case, the quantization matrix used for subblocks using intra-inter-mixed prediction and the quantization matrix used for subblocks using intra-inter-combined prediction can be made common. This makes it possible to apply quantization using a quantization matrix with the same quantization control properties to subblocks that use predictions that share the common characteristic of using both predicted pixels by intra-prediction and predicted pixels by inter-prediction within the same subblock. Moreover, it is possible to reduce the amount of code corresponding to the quantization matrix for the new prediction method.

[0093] Furthermore, while the first embodiment used an input image as the data to be encoded, the data to be encoded is not limited to images. For example, a two-dimensional data array, which is feature data used in machine learning such as object recognition, can be encoded in the same way as an input image to generate a bitstream and output. This makes it possible to efficiently encode feature data used in machine learning.

[0094] [Second Embodiment] The image decoding device according to this embodiment decodes the quantization coefficients for the block to be decoded from the bitstream, derives conversion coefficients from the quantization coefficients using a quantization matrix, and derives the prediction error for the block to be decoded by inverse frequency transformation of the conversion coefficients. The image decoding device then generates a prediction image by applying an intra-predicted image obtained by intra-prediction to a portion of the block to be decoded, and an inter-predicted image obtained by inter-prediction to other portions of the block that differ from the portion, and decodes the block to be decoded using the generated prediction image and the prediction error.

[0095] This embodiment describes an image decoding device that decodes a bitstream encoded by an image encoding device according to the first embodiment. First, an example of the functional configuration of the image decoding device according to this embodiment will be described using the block diagram in Figure 2.

[0096] The control unit 250 controls the operation of the entire image decoding device. The separation and decoding unit 202 acquires the bitstream encoded by the image encoding device according to the first embodiment. The method of acquiring the bitstream is not limited to a specific method. For example, the bitstream output from the image encoding device according to the first embodiment may be acquired via a network, or it may be acquired from a memory where the bitstream has been temporarily stored. The separation and decoding unit 202 then separates the acquired bitstream into encoded data related to the decoding process and coefficients, and decodes the encoded data present in the header portion of the bitstream. In this embodiment, the separation and decoding unit 202 separates the encoded data of the quantization matrix from the bitstream and supplies the encoded data to the decoding unit 209. The separation and decoding unit 202 also separates the encoded data of the input image from the bitstream and supplies the encoded data to the decoding unit 203. In other words, the separation and decoding unit 202 operates in the reverse direction of the integrated encoding unit 111 in Figure 1.

[0097] The decoding unit 209 decodes the encoded data supplied from the separation decoding unit 202 to reconstruct the quantization matrix. The decoding unit 203 decodes the encoded data supplied from the separation decoding unit 202 to reconstruct the quantization coefficients and prediction information.

[0098] The inverse quantization / inverse transformation unit 204 operates in the same manner as the inverse quantization / inverse transformation unit 106 in the image coding device according to the first embodiment. The inverse quantization / inverse transformation unit 204 selects a quantization matrix corresponding to the prediction corresponding to the quantization coefficient to be decoded from among the quantization matrices decoded by the decoding unit 209, and uses the selected quantization matrix to inverse quantize the quantization coefficient to reconstruct the orthogonal transformation coefficient. The inverse quantization / inverse transformation unit 204 then reconstructs the prediction error by performing an inverse orthogonal transformation on the reconstructed orthogonal transformation coefficient.

[0099] The image playback unit 205 generates a predicted image by referring to the image stored in the frame memory 206 based on the predicted information decoded by the decoding unit 203. The image playback unit 205 then generates a reproduced image by adding the prediction error obtained by the inverse quantization / inverse transform unit 204 to the generated predicted image, and stores the generated reproduced image in the frame memory 206.

[0100] The in-loop filter unit 207 performs in-loop filtering, such as deblocking filtering and sample adaptive offsetting, on the playback image stored in the frame memory 206. The playback image stored in the frame memory 206 is output as appropriate by the control unit 250. The output destination of the playback image is not limited to a specific destination; for example, the playback image may be displayed on the display screen of a display device such as a monitor, or the playback image may be output to a projection device such as a projector.

[0101] Next, the operation of the image decoding device having the above configuration (bitstream decoding process) will be described. The separation and decoding unit 202 acquires the bitstream generated by the image encoding device, separates encoded data related to the decoding process and coefficients from the bitstream, and decodes the encoded data present in the bitstream header. The separation and decoding unit 202 extracts encoded data of the quantization matrix from the bitstream sequence header in Figure 6(a) and supplies the extracted encoded data to the decoding unit 209. The separation and decoding unit 202 also supplies encoded data of the picture data in subblock units to the decoding unit 203.

[0102] The decoding unit 209 decodes the encoded data of the quantization matrix supplied from the separation decoding unit 202 to reconstruct a one-dimensional array. More specifically, the decoding unit 209 refers to the encoding tables illustrated in Figures 11(a) and 11(b) to decode the binary codes in the encoded data of the quantization matrix into difference values ​​and generate a one-dimensional array by arranging them. For example, when the decoding unit 209 decodes the encoded data of the quantization matrices in Figures 8(a) to 8(c), the one-dimensional arrays in Figures 10(a) to 8(c) are reconstructed, respectively. In this embodiment, as in the first embodiment, decoding is performed using the encoding table shown in Figure 11(a) (or Figure 11(b)), but the encoding table is not limited to this, and other encoding tables may be used as long as they are the same as those in the first embodiment.

[0103] Furthermore, the decoding unit 209 reproduces each element value of the quantization matrix from each difference value of the reproduced one-dimensional array. This is to perform a process reverse to the process performed by the encoding unit 113 to generate a one-dimensional array from the quantization matrix. That is, the value of the element at the beginning of the one-dimensional array is the element value at the upper left corner of the quantization matrix. The value obtained by adding the value of the element at the beginning of the one-dimensional array to the value of the second element from the beginning of the one-dimensional array is the second element value in the above "prescribed order". The value obtained by adding the value of the (n - 1)-th element from the beginning of the one-dimensional array to the value of the n-th element (2 < n ≦ N: N is the number of elements of the one-dimensional array) from the beginning of the one-dimensional array is the n-th element value in the above "prescribed order". For example, the decoding unit 209 reproduces the quantization matrices of FIGS. 8(a) to (c) from the one-dimensional arrays of FIGS. 10(a) to (c) using the order shown in FIG. 9, respectively.

[0104] The decoding unit 203 decodes the quantization coefficients and prediction information by decoding the encoded data of the input image supplied from the separation decoding unit 202.

[0105] The inverse quantization / inverse transformation unit 204 specifies the "prediction mode corresponding to the quantization coefficient to be decoded" included in the prediction information decoded by the decoding unit 203, and selects the quantization matrix corresponding to the specified prediction mode from the quantization matrices reproduced by the decoding unit 209. Then, the inverse quantization / inverse transformation unit 204 inverse-quantizes the quantization coefficient using the selected quantization matrix to reproduce the orthogonal transformation coefficients. Then, the inverse quantization / inverse transformation unit 204 performs an inverse orthogonal transformation on the reproduced orthogonal transformation coefficients to reproduce the prediction error, and supplies the reproduced prediction error to the image reproduction unit 205.

[0106] The image playback unit 205 generates a predicted image by referring to the image stored in the frame memory 206 based on the predicted information decoded by the decoding unit 203. In this embodiment, three types of predictions are used, intra prediction, inter prediction, and intra / inter mixed prediction, similar to the prediction unit 104 in the first embodiment. The specific prediction process is the same as that of the prediction unit 104 described in the first embodiment, so the explanation is omitted. The image playback unit 205 then generates a reconstructed image by adding the prediction error obtained by the inverse quantization / inverse transform unit 204 to the generated predicted image, and stores the generated reconstructed image in the frame memory 206. The reconstructed image stored in the frame memory 206 becomes a prediction reference candidate that is referenced when decoding other subblocks.

[0107] The in-loop filter unit 207 operates similarly to the in-loop filter unit 109 described above, performing in-loop filtering such as deblocking filtering and sample adaptive offsetting on the playback image stored in the frame memory 206. The playback image stored in the frame memory 206 is output as appropriate by the control unit 250.

[0108] The decoding process in the image decoding device according to this embodiment will be explained with reference to the flowchart in Figure 4. In step S401, the separation decoding unit 202 acquires the encoded bitstream. The separation decoding unit 202 then separates the encoded data of the quantization matrix from the acquired bitstream and supplies the encoded data to the decoding unit 209. The separation decoding unit 202 also separates the encoded data of the input image from the bitstream and supplies the encoded data to the decoding unit 203.

[0109] In step S402, the decoding unit 209 decodes the encoded data supplied from the separation decoding unit 202 to reconstruct the quantization matrix. In step S403, the decoding unit 203 decodes the encoded data supplied from the separation decoding unit 202 to reconstruct the quantization coefficients and prediction information of the subblock to be decoded.

[0110] In step S404, the inverse quantization / inverse transformation unit 204 identifies the "prediction mode corresponding to the quantization coefficient of the subblock to be decoded" contained in the prediction information decoded by the decoding unit 203. The inverse quantization / inverse transformation unit 204 then selects the quantization matrix corresponding to the identified prediction mode from the quantization matrices reconstructed by the decoding unit 209. For example, if the prediction mode identified for the subblock to be decoded is intra-prediction, the intra-prediction quantization matrix in Figure 8(a) to (c) is selected from the quantization matrices in Figures 8(a) to (c). If the prediction mode identified for the subblock to be decoded is inter-prediction, the inter-prediction quantization matrix in Figure 8(b) is selected. If the prediction mode identified for the subblock to be decoded is a mixed intra-inter-prediction, the intra-inter-inter-prediction quantization matrix in Figure 8(c) is selected. The inverse quantization / inverse transformation unit 204 then uses the selected quantization matrix to inverse quantize the quantization coefficient of the subblock to be decoded and reconstructs the orthogonal transformation coefficients. The inverse quantization / inverse transformation unit 204 then performs an inverse orthogonal transformation on the reconstructed orthogonal transformation coefficients to reconstruct the prediction error of the subblock to be decoded, and supplies the reconstructed prediction error to the image reconstruction unit 205.

[0111] In step S405, the image playback unit 205 generates a predicted image of the subblock to be decoded by referring to the image stored in the frame memory 206 based on the predicted information decoded by the decoding unit 203. The image playback unit 205 then generates a reconstructed image of the subblock to be decoded by adding the prediction error of the subblock to be decoded obtained by the inverse quantization / inverse transform unit 204 to the generated predicted image, and stores the generated reconstructed image in the frame memory 206.

[0112] In step S406, the control unit 250 determines whether the processing in steps S403 to S405 has been performed for all subblocks. If the result of this determination is that the processing in steps S403 to S405 has been performed for all subblocks, the process proceeds to step S407. On the other hand, if there are still subblocks that have not been processed in steps S403 to S405, the process proceeds to step S403 in order to perform the processing in steps S403 to S405 for those subblocks.

[0113] In step S407, the in-loop filter unit 207 performs in-loop filtering, such as deblocking filtering and sample adaptive offsetting, on the replayed image generated in step S405 and stored in the frame memory 206.

[0114] Through this process, even for subblocks generated in the first embodiment using intra-inter-mixed prediction, it is possible to decode a bitstream with improved image quality by controlling quantization for each frequency component.

[0115] <Variation> In the second embodiment, separate quantization matrices were prepared for intra-prediction, inter-prediction, and intra-inter-mixed prediction, and the quantization matrices corresponding to each prediction were decoded. However, some of these matrices may be shared.

[0116] For example, to dequantize the quantization coefficients of the orthogonal transformation coefficients of the prediction error obtained based on intra-internal mixed prediction, one may use the quantization matrix corresponding to intra-prediction instead of the quantization matrix corresponding to intra-internal mixed prediction. That is, for example, to dequantize the quantization coefficients of the orthogonal transformation coefficients of the prediction error obtained based on intra-internal mixed prediction, one may use the quantization matrix for intra-prediction shown in Figure 8(a). In this case, decoding of the quantization matrix corresponding to intra-internal mixed prediction can be omitted. In other words, it is possible to decode a bitstream with a reduced amount of encoded data for the quantization matrix included in the bitstream, and to obtain a decoded image with reduced image quality degradation caused by errors due to intra-prediction, such as block distortion.

[0117] Furthermore, in order to dequantize the quantization coefficients of the orthogonal transformation coefficients of the prediction error obtained based on intra-internal mixed prediction, it is also possible to decode and use the quantization matrix corresponding to inter-prediction instead of the quantization matrix corresponding to intra-internal mixed prediction. That is, for example, in order to dequantize the quantization coefficients of the orthogonal transformation coefficients of the prediction error obtained based on intra-internal mixed prediction, the quantization matrix for inter-prediction shown in Figure 8(b) may be used. In this case, decoding of the quantization matrix corresponding to intra-internal mixed prediction can be omitted. In other words, it is possible to decode a bitstream with a reduced amount of encoded data for the quantization matrix included in the bitstream, and to obtain a decoded image with reduced image quality degradation caused by errors due to inter-prediction, such as jerky motion.

[0118] Furthermore, in the predicted image of a subblock where intra-inter prediction has been performed, the quantization matrix used for inverse quantization of the subblock may be determined according to the respective sizes of the region of "predicted pixels obtained by intra prediction" and the region of "predicted pixels obtained by inter prediction".

[0119] For example, suppose that subblock 1200 is divided into divided region 1200c and divided region 1200d, as shown in Figure 12(e). Then, suppose that the predicted pixels for divided region 1200c are determined by "predicted pixels obtained by intra-prediction", and the predicted pixels for divided region 1200d are determined by "predicted pixels obtained by inter-prediction". Also, suppose that the size (area (number of pixels)) S1 of divided region 1200c and the size (area (number of pixels)) S2 of divided region 1200d are in a ratio of 1:3.

[0120] In this case, the size of the partitioned region 1200d to which interpretation is applied is larger than the size of the partitioned region 1200c to which intrapretation is applied in subblock 1200. Therefore, the inverse quantization / inverse transformation unit 204 applies a quantization matrix corresponding to interpretation to the inverse quantization of the quantization coefficients of this subblock 1200.

[0121] Furthermore, if the size of the partitioned region 1200d to which interpretation is applied is smaller than the size of the partitioned region 1200c to which intraprediction is applied in subblock 1200, the inverse quantization / inverse transformation unit 204 applies a quantization matrix corresponding to the intraprediction to the inverse quantization of the quantization coefficients of subblock 1200.

[0122] This reduces image quality degradation in larger segmented regions while omitting the decoding of the quantization matrix corresponding to intra-inter-mixed prediction. Therefore, it becomes possible to decode a bitstream with a reduced amount of encoded data for the quantization matrix included in the bitstream.

[0123] Alternatively, a quantization matrix may be generated as the quantization matrix corresponding to the intra-internal mixed prediction by combining the "quantization matrix corresponding to intra-prediction" and the "quantization matrix corresponding to inter-prediction" according to the ratio of S1 and S2. For example, the inverse quantization / inverse transformation unit 204 may generate the quantization matrix corresponding to the intra-internal mixed prediction using the above equation (1).

[0124] Thus, since the quantization matrix corresponding to the mixed intra- and inter-prediction can be generated as needed, the encoding of the quantization matrix can be omitted. Therefore, it is possible to decode a bitstream with a reduced amount of encoded data for the quantization matrix included in the bitstream. Furthermore, it is possible to decode a bitstream with improved image quality by performing appropriate quantization control according to the ratio of the sizes of the regions in which intra-prediction and inter-prediction are used.

[0125] Furthermore, in the second embodiment, the quantization matrix applied to the subblock to which intra-inter-mixed prediction is applied is uniquely determined, but as in the first embodiment, it may be configured to be selectable by introducing an identifier. This makes it possible to decode a bitstream that selectively reduces the amount of encoded data of the quantization matrix included in the bitstream and enables unique quantization control for the subblock to which intra-inter-mixed prediction is applied.

[0126] Furthermore, in the second embodiment, a prediction image is decoded that includes prediction pixels for one of the divided regions of a subblock (first prediction pixels) and prediction pixels for the other divided region (second prediction pixels). However, the prediction image to be decoded is not limited to such a prediction image. For example, similar to the modification of the first embodiment, the prediction image may be one in which the third prediction pixel, calculated by a weighted average of the first and second prediction pixels included in the region near the boundary between one divided region and the other (boundary region), is the prediction pixel of the boundary region. In this case, as in the first embodiment, the prediction pixel value of the corresponding region corresponding to one of the divided regions in the prediction image becomes the first prediction pixel, and the prediction pixel value of the corresponding region corresponding to the other divided region in the prediction image becomes the second prediction pixel. The prediction pixel value of the corresponding region corresponding to the boundary region in the prediction image becomes the third prediction pixel. This makes it possible to suppress the degradation of image quality in the boundary region of divided regions in which different predictions are used, and to decode a bitstream with improved image quality.

[0127] Furthermore, in the second embodiment, three types of predictions were explained as examples: intra-prediction, inter-prediction, and intra-inter-mixed prediction. However, the types and number of predictions are not limited to these examples. For example, intra-inter-combined prediction (CIIP), which is used in VVC, may also be used. In this case, the quantization matrix used for subblocks using intra-inter-mixed prediction and the quantization matrix used for subblocks using intra-inter-combined prediction can be made common. This makes it possible to decode bitstreams to which quantization by a quantization matrix with the same quantization control properties is applied to subblocks that use prediction methods that share the commonality of using both predicted pixels by intra-prediction and predicted pixels by inter-prediction within the same subblock. Moreover, it is also possible to decode bitstreams with a reduced code amount by the amount of the quantization matrix corresponding to the new prediction method.

[0128] Furthermore, although the second embodiment was described as decoding an input image to be encoded from a bitstream, the object to be decoded is not limited to an image. For example, the configuration may be such that a two-dimensional data array, which is feature data used in machine learning such as object recognition, is encoded in the same way as the input image, and the two-dimensional data array is decoded from the bitstream containing the encoded data. This makes it possible to decode a bitstream that has been efficiently encoded with feature data used in machine learning.

[0129] [Third Embodiment] The image encoding device according to this embodiment encodes an image by performing prediction processing on a block-by-block basis. The image encoding device determines the intensity of the deblocking filter processing performed on the boundary between a first block in the image and a second block adjacent to the first block in the image, and performs the deblocking filter processing on the boundary according to the determined intensity. As the prediction processing, one of the following is used: a first prediction mode (intra prediction) that derives the predicted pixels of the block to be encoded using pixels in the image containing the block to be encoded; a second prediction mode (inter prediction) that derives the predicted pixels of the block to be encoded using pixels in another image different from the image containing the block to be encoded; and a third prediction mode (intra-inter mixed prediction) that derives the predicted pixels for a part of the block to be encoded using pixels in the image containing the block to be encoded, and derives the predicted pixels for other parts of the block to be encoded using pixels in another image different from the image containing the block to be encoded. Furthermore, in determining the intensity of the deblocking filter, the image encoding device determines the above intensity as the first intensity if at least one of the first block and the second block is a block to which the first prediction mode is applied. Also, if at least one of the first block and the second block is a block to which the third prediction mode is applied, the image encoding device determines the above intensity as the intensity based on the above first intensity.

[0130] An example of the functional configuration of the image encoding device according to this embodiment will be explained using the block diagram in Figure 13. In Figure 13, the same reference numbers are used for the same functional parts as those shown in Figure 1, and the explanation of these functional parts will be omitted.

[0131] In this embodiment, the transformation / quantization unit 105 is described as using a predetermined quantization matrix when quantizing the orthogonal transformation coefficients. However, as in the above embodiment, a quantization matrix corresponding to intra-inter mixed prediction may also be used.

[0132] The in-loop filter unit 1309 performs in-loop filtering, such as deblocking filtering, on the image (subblock) stored in the frame memory 108, according to the filter intensity (bS value) determined by the determination unit 1313. The in-loop filter unit 1309 then stores the in-loop filtered image in the frame memory 108.

[0133] The determination unit 1313 determines the filter strength (bS value) of the deblocking filter process performed on the boundary between two adjacent subblocks. More specifically, the determination unit 1313 determines the bS value, which is the filter strength of the deblocking filter process performed on the boundary between subblock P and the subblock Q adjacent to subblock P, based on the conditions (Condition 1) to (Condition 6) below that are satisfied.

[0134] (Condition 1) If both subblock P and subblock Q are subblocks of the "BDPCM mode that directly encodes the difference values ​​of adjacent pixels without performing any transformation on the predicted difference image that has undergone intra-prediction in the horizontal or vertical direction," then the bS value is set to 0. (Condition 2) If at least one of subblocks P and Q is a subblock that has undergone intra prediction or intra / inter mixed prediction, the bS value shall be 2. (Condition 3) If the boundary between subblock P and subblock Q is the boundary of a subblock that is the unit of transformation, and the orthogonal transformation coefficients of at least one of the subblocks include non-zero orthogonal transformation coefficients, then the bS value is set to 1. (Condition 4) If the reference images for motion compensation are different between subblock P and subblock Q, or if the number of motion vectors is different, the bS value should be set to 1. (Condition 5) If the absolute value of the difference between the motion vector in subblock P and the motion vector in subblock Q is 0.5 pixels or more, the bS value is set to 1. (Condition 6) If not (Condition 1) to (Condition 5), the bS value is set to 0. Here, deblocking filtering is not performed on subblock boundaries (edges) where the bS value is 0. For subblock boundaries (edges) where the bS value is 1 or greater, the deblocking filter is determined based on the gradient and activity near the subblock boundary. In principle, the larger the bS value, the stronger the deblocking filter applied.

[0135] In this embodiment, deblocking filtering is not performed on subblock boundaries where the bS value is 0, and is performed on subblock boundaries where the bS value is 1 or greater, but this is not limited to this. For example, there may be more or fewer types of intensities for the deblocking filtering.

[0136] Furthermore, the processing method may differ depending on the intensity of the deblocking filter. For example, the bS value may take five values ​​from 0 to 4, as in the deblocking filter of H.264.

[0137] Furthermore, in this embodiment, the bS value of the deblocking filter applied to the subblock boundary using intra-inter-mixing prediction is the same as the bS value of the deblocking filter applied to the subblock boundary using intra-prediction. This bS value is 2, which indicates the maximum filter strength, but it is not limited to this. For example, an intermediate bS value may be provided between bS value = 1 (an example of another bS value) and bS value = 2 in this embodiment, and this intermediate bS value may be used if at least one of subblocks P and Q is a subblock that has undergone intra-inter-mixing prediction. In that case, it is possible to apply the same deblocking filter application to the luminance component as when the normal bS value is 2, and to apply a deblocking filter application with a weaker correction strength than when the bS value is 2 to the chromatic difference component. This makes it possible to apply a deblocking filter application with an intermediate correction strength to the subblock boundary using intra-inter-mixing prediction. Alternatively, the bS value may be determined for each boundary pixel based on whether each pixel within a subblock using intra-internal mixed prediction is predicted by intra-prediction or inter-prediction. In this case, the bS value for pixels predicted by intra-prediction will always be 2, and the bS value for pixels predicted by inter-prediction can be determined based on the above conditions (1) to (6).

[0138] Furthermore, while the bS value is used as the filter strength in this embodiment, it is not limited to this. For example, another variable may be defined as the filter strength instead of the bS value, or the coefficients or filter length of the deblocking filter may be changed directly.

[0139] Next, the deblocking filter processing in the in-loop filter section 1309 of this embodiment will be described in more detail. The deblocking filter processing is performed on the boundaries of subblocks that form the units of prediction processing or transformation processing. The filter length of the deblocking filter depends on the size (number of pixels) of the subblock, and if the subblock size is 32 pixels or more, it is applied up to 7 pixels from the boundary. Similarly, if the subblock size is 4 pixels or less, only the pixel value of the 1-pixel line adjacent to the boundary is updated.

[0140] In this embodiment, all subblocks are assumed to be 8x8 pixels in size and undergo prediction and transformation processing, but this is not limited to this, and the size of the subblock performing prediction and the size of the subblock performing transformation processing may be different. For example, as in Subblock Transform (SBT) in VVC, the subblock to which the transformation processing is applied may be a subblock that further divides the subblock that performs prediction processing. Alternatively, it may be larger, such as 32x32 pixels, or not be a square, such as 16x8 pixels.

[0141] In Figure 14, subblocks P and Q are adjacent 8x8 pixel subblocks separated by a boundary, and also serve as units for orthogonal transformations. p00~p33 represent pixels (pixel values) belonging to subblock P, and q00~q33 represent pixels (pixel values) belonging to subblock Q. The pixel groups p00~p33 and q00~q33 are adjacent across the boundary. First, if the bS value is 1 or greater with respect to brightness, the in-loop filter unit 1309 determines whether or not to perform deblocking filtering on the boundary between subblock P and subblock Q, for example, according to the following formula.

[0142] |p20-2×p10+p00|+|p23-2×p13+p03|+|q20-2×q10+q00|+|q23-2×q13+q03|<β Here, β is a value corresponding to the average of the quantization step value in subblock P and the quantization step value in subblock Q. For example, among the various β values ​​registered in the table, the β registered in association with the average value is obtained. Only if this equation is satisfied does the in-loop filter unit 1309 determine to perform deblocking filtering on the boundary between subblock P and subblock Q.

[0143] If it is determined that deblocking filtering should be performed, the in-loop filter unit 1309 determines whether to use a strong filter or a weak filter, which have different smoothing effects. For example, the in-loop filter unit 1309 determines to use a strong filter if all of the following six equations ((1) to (6)) are satisfied. On the other hand, the in-loop filter unit 1309 determines to use a weak filter if even one of the following six equations ((1) to (6)) is not satisfied. (1) 2×(|p20-2×p10+p00|+|q20-2×q10+q00|)<(β>>2) (2) 2×(|p23-2×p13+p03|+|q23-2×q13+q03|)<(β>>2) (3) |p30-p00|+|q00-q30|<(β>>3) (4) |p33-p03|+|q03-q33|<(β>>3) (5) |p00-q00|<((5×tc+1)>>1) (6) |p03-q03|<((5×tc+1)>>1) Here, >>N(N=1~3) means an N-bit arithmetic right shift operation, and tc is a parameter that determines the maximum amount of pixel value correction. tc can be obtained, for example, by the following process: That is, the average value qP of the quantization step value in subblock P, the quantization step value in subblock Q, and the bS value is given by the following formula qP = qP + 2x(bS - 1) The correction is performed accordingly, and the tc registered in the table that is associated with the corrected qP value is retrieved. From this formula, it can be seen that when the bS value is 2, the corrected qP value becomes larger. The table is set up so that the tc value increases as the qP value increases, so the larger the bS value, the stronger the deblocking filter applied, which corrects the pixel value.

[0144] Strong filtering for luminance (filtering using a strong filter) can be expressed by the following equation, where p'0k, p'1k, and p'2k are the pixel values ​​after deblocking filtering in subblock P, and q'0k, q'1k, and q'2k are the pixel values ​​after deblocking filtering in subblock Q (k=0 to 3).

[0145] p'0k=Clip3(p0k-3×tc,p0k+3×tc,(p2k+2×p1k+2×p0k+2×q0k+q1k+4)>>3) p'1k=Clip3(p1k-2×tc,p1k+2×tc,(p2k+p1k+p0k+q0k+2)>>2) p'2k=Clip3(p2k-1×tc,p2k+1×tc,(2×p3k+3×p2k+p1k+p0k+q0k+4)>>3) q'0k=Clip3(q0k-3×tc,q0k+3×tc,(q2k+2×q1k+2×q0k+2×p0k+p1k+4)>>3) q'1k=Clip3(q1k-2×tc,q1k+2×tc,(q2k+q1k+q0k+p0k+2)>>2) q'2k=Clip3(q2k-1×tc,q2k+1×tc,(2×q3k+3×q2k+q1k+q0k+p0k+4)>>3) Here, Clip3(a,b,c) is a function that performs clipping so that the range of c is a≦c≦b. Also, weak filtering with respect to brightness (filtering using a weak filter) is expressed by the following formula.

[0146] Δ=(9×(q0k-p0k)-3×(q1k-p1k)+8)>>4 |Δ|<10×tc If the above conditions are not met, the deblocking filter process will not be performed. If the conditions are met, the following processing will be performed on p0k and q0k according to the formula below.

[0147] Δ=Clip3(-tc,tc,Δ) p'0k=Clip1(p0k+Δ) q'0k = Clip1(q0k - Δ) Here, Clip1(a) is a function that performs clipping so that the range of a is 0 ≤ a ≤ (the maximum value that can be represented by the bit depth of the luminance or chrominance signal). For example, if the luminance is 8 bits (the maximum value that can be represented by the bit depth of the luminance), it will be 255, and if the luminance is 10 bits (the maximum value that can be represented by the bit depth of the luminance), it will be 1023.

[0148] Furthermore, the following conditions |p20-2×p10+p00|+|p23-2×p13+p03|<(β+(β>>1))>>3) |q20-2×q10+q00|+|q23-2×q13+q03|<(β+(β>>1))>>3) When the following condition is met, a deblocking filter is performed on p1k and q1k according to the following formula.

[0149] Δp=Clip3(-(tc>>1),tc>>1,(((p2k+p0k+1)>>1-p1k+Δ)>>1) p'1k=Clip1(p1k+Δp) Δq=Clip3(-(tc>>1),tc>>1,(((q2k+q0k+1)>>1-q1k+Δ)>>1) q'1k = Clip1(q1k + Δq) In deblocking filtering for color difference, when the filter length is 1, deblocking filtering is performed according to the following formula only when the bS value is 2.

[0150] Δ=Clip3(-tc,tc,(((q0k-p0k)<<2)+p1k-q1k+4)>>3)) p'0k=Clip1(p0k+Δ) q'0k = Clip1(q0k - Δ) In this embodiment, the deblocking filter processing for color difference is performed only when the bS value is 2 when the filter length is 1, but it is not limited to this. For example, the deblocking filter processing may be performed when the bS value is not 0, or it may be performed only when the filter length is longer than 1.

[0151] In this embodiment, the bS value was used to determine whether or not to apply a deblocking filter and to calculate the maximum amount of pixel value correction by the deblocking filter. Separately, a strong filter with a high smoothing effect and a weak filter with a low smoothing effect were used depending on the pixel value conditions. However, this is not the only way. For example, the filter length may be determined according to the bS value, or only the strength of the smoothing effect may be determined by the bS value.

[0152] The encoding process performed by the image encoding device described above will now be explained according to the flowchart in Figure 15. Note that the process according to the flowchart in Figure 15 is for encoding a single input image. Therefore, when encoding images for each frame of a video, or multiple images captured periodically or irregularly, the process according to the flowchart in Figure 15 will be repeated for each image.

[0153] In step S1501, the integrated encoding unit 111 encodes various header information necessary for encoding the input image to generate header code data.

[0154] In step S1502, the division unit 102 divides the input image into multiple basic blocks and outputs each of the divided basic blocks. Then, the prediction unit 104 divides each basic block into multiple subblocks.

[0155] In step S1503, the prediction unit 104 selects one of the unselected subblocks in the input image as the selected subblock and determines the prediction mode for the selected subblock. The prediction unit 104 then performs a prediction on the selected subblock according to the determined prediction mode and obtains the predicted image, prediction error, and prediction information for the selected subblock.

[0156] In step S1504, the transformation / quantization unit 105 applies an orthogonal transformation process to the prediction error of the selected subblock obtained in step S1503 to generate orthogonal transformation coefficients, and then quantizes these orthogonal transformation coefficients using a quantization matrix to obtain quantization coefficients.

[0157] In step S1505, the inverse quantization / inverse transformation unit 106 generates orthogonal transformation coefficients by performing inverse quantization on the quantization coefficients of the selected subblock obtained in step S1504 using the above-mentioned quantization matrix. The inverse quantization / inverse transformation unit 106 then performs an inverse orthogonal transformation on the generated orthogonal transformation coefficients to generate (reconstruct) the prediction error.

[0158] In step S1506, the image playback unit 107 generates a predicted image from the images stored in the frame memory 108 based on the prediction information acquired in step S1503, and plays back the image of the subblock by adding the predicted image to the prediction error generated in step S1505. The image playback unit 107 then stores the played-back image in the frame memory 108.

[0159] In step S1507, the coding unit 110 entropy encodes the quantization coefficients obtained in step S1504 and the prediction information obtained in step S1503 to generate coded data.

[0160] The integrated encoding unit 111 then generates a bitstream by multiplexing the header code data generated in step S1501 and the encoded data generated by the encoding unit 110 in step S1507.

[0161] In step S1508, the control unit 150 determines whether all subblocks in the input image have been selected as selected subblocks. If, as a result of this determination, all subblocks in the input image have been selected as selected subblocks, the process proceeds to step S1509. On the other hand, if there is one or more subblocks in the input image that have not yet been selected as selected subblocks, the process proceeds to step S1503.

[0162] In step S1509, the determination unit 1313 determines a bS value for each boundary between adjacent subblocks. Then, in step S1510, the in-loop filter unit 1309 performs in-loop filtering on the image stored in the frame memory 108, such as a deblocking filter with a filter intensity based on the bS value determined in step S1509. More specifically, the in-loop filter unit 1309 performs in-loop filtering on the boundaries of subblocks that are the units of orthogonal transformation of the image stored in the frame memory 108, such as a deblocking filter with a filter intensity based on the bS value determined by the determination unit 1313 for each boundary. Then, the in-loop filter unit 109 stores the in-loop filtered image in the frame memory 108.

[0163] Thus, according to this embodiment, a deblocking filter with a high distortion correction effect can be set for the boundaries of subblocks using intra-inter-mixed prediction, thereby suppressing block distortion and improving image quality. Furthermore, since no new calculations are required to calculate the filter strength, the complexity of the implementation is not increased.

[0164] Furthermore, in this embodiment, the filter strength was used to determine whether or not to apply a deblocking filter and to calculate the maximum amount of pixel value correction by the deblocking filter. However, it is not limited to this. For example, when the filter strength is greater, a deblocking filter with a longer tap length and a higher correction effect can be used, and when the filter strength is smaller, a deblocking filter with a shorter tap length and a lower correction effect can be used.

[0165] Furthermore, although this embodiment assumes that only three types of prediction are used—intra-prediction, inter-prediction, and intra-inter-mixed prediction—it is not limited to these, and for example, intra-inter-composite prediction (CIIP) used in VVC may also be used. In this case, the bS value used for a subblock using intra-inter-mixed prediction can be the same as the bS value used when intra-inter-composite prediction is used. This makes it possible to encode bitstreams by applying a deblocking filter of the same strength to subblocks that use predictions that share the common characteristic of using both intra-prediction pixels and inter-prediction pixels within the same subblock.

[0166] [Fourth Embodiment] The image decoding device according to this embodiment is an image decoding device that decodes an encoded image block by block. Such an image decoding device decodes an image by performing prediction processing for each block. It also determines the intensity of the deblocking filter processing to be performed on the boundary between a first block and a second block adjacent to the first block, and performs the deblocking filter processing on the boundary according to the determined intensity. Furthermore, the image decoding device uses one of the following prediction modes as the prediction processing: a first prediction mode (intra prediction) that derives the predicted pixels of the block to be decoded using pixels in the image containing the block to be decoded; a second prediction mode (inter prediction) that derives the predicted pixels of the block to be decoded using pixels in another image different from the image containing the block to be decoded; and a third prediction mode (intra-inter mixed prediction) that derives the predicted pixels for a part of the block to be decoded using pixels in the image containing the block to be decoded, and derives the predicted pixels for other parts of the block that are different from the part of the block to be decoded using pixels in another image different from the image containing the block to be decoded. Here, in determining the intensity of the deblocking filter, if at least one of the first block and the second block is a block to which the first prediction mode is applied, the intensity is determined to be the first intensity. Furthermore, if at least one of the first block and the second block is a block to which the third prediction mode is applied, the intensity is determined to be the intensity based on the first intensity.

[0167] This embodiment describes an image decoding device that decodes a bitstream encoded by an image encoding device according to the third embodiment. First, an example of the functional configuration of the image decoding device according to this embodiment will be described using the block diagram in Figure 16. In Figure 16, the same reference numerals are used for the same functional parts as shown in Figure 2, and the description of these functional parts will be omitted.

[0168] The in-loop filter unit 1607 performs in-loop filtering, such as deblocking filtering, on the subblock boundaries of the replayed image stored in the frame memory 206, with a filter intensity corresponding to the bS value determined by the determination unit 1609 for the subblock boundary. The determination unit 1609 determines the filter intensity (bS value) for deblocking filtering performed on the boundary between two adjacent subblocks, in the same manner as the determination unit 1313.

[0169] In this embodiment, as in the third embodiment, the bS value was used to determine whether or not to apply a deblocking filter and to calculate the maximum amount of pixel value correction by the deblocking filter. Separately, a strong filter with a high smoothing effect and a weak filter with a low smoothing effect were used depending on the pixel value conditions. However, this is not limited to this. For example, the filter length may be determined according to the bS value, or only the strength of the smoothing effect may be determined by the bS value.

[0170] The decoding process in the image decoding device according to this embodiment will be explained with reference to the flowchart in Figure 17. In step S1701, the separation decoding unit 202 acquires the bitstream. The separation decoding unit 202 then separates the encoded data of the input image from the bitstream and supplies the encoded data to the decoding unit 203, while also decoding the header encoded data in the bitstream.

[0171] In step S1702, the decoding unit 203 decodes the encoded data supplied from the separation decoding unit 202 and reconstructs the quantization coefficients and prediction information of the subblock to be decoded.

[0172] In step S1703, the inverse quantization / inverse transformation unit 204 uses the quantization matrix to inverse quantize the quantization coefficients of the subblock to be decoded and reconstructs the orthogonal transformation coefficients. The inverse quantization / inverse transformation unit 204 then performs an inverse orthogonal transformation on the reconstructed orthogonal transformation coefficients to reconstruct the prediction error of the subblock to be decoded, and supplies the reconstructed prediction error to the image regeneration unit 205.

[0173] In step S1704, the image playback unit 205 generates a predicted image of the subblock to be decoded by referring to the image stored in the frame memory 206 based on the predicted information decoded by the decoding unit 203. The image playback unit 205 then generates a reconstructed image of the subblock to be decoded by adding the prediction error obtained by the inverse quantization / inverse transform unit 204 to the generated predicted image, and stores the generated reconstructed image in the frame memory 206.

[0174] In step S1705, the control unit 250 determines whether the processing in steps S1702 to S1704 has been performed for all subblocks. If the result of this determination is that the processing in steps S1702 to S174 has been performed for all subblocks, the process proceeds to step S1706. On the other hand, if there are still subblocks that have not been processed in steps S1702 to S1704, the process proceeds to step S1702 in order to perform the processing in steps S1702 to S1704 for those subblocks.

[0175] In step S1706, the determination unit 1609 determines the filter strength of the deblocking filter process performed on the boundary between two adjacent subblocks, similar to the determination unit 1313 described in the third embodiment. The type of prediction applied to each subblock (intra prediction, inter prediction, intra / inter mixed prediction) is recorded in the prediction information, so the prediction applied to each subblock can be identified by referring to this prediction information.

[0176] In step S1707, the in-loop filter unit 1607 performs in-loop filtering on the sub-block boundaries of the replayed image stored in the frame memory 206, such as a deblocking filter with a filter strength corresponding to the bS value determined by the determination unit 1609 for the sub-block boundaries.

[0177] Thus, according to this embodiment, an appropriate deblocking filter can be applied when decoding a bitstream containing "subblocks encoded with intra-inter-mixed prediction" generated by the image encoding device according to the third embodiment.

[0178] [Fifth Embodiment] Each of the functional units shown in Figures 1, 2, 13, and 16 may be implemented in hardware, or each functional unit except for the holding unit 103 and frame memories 108 and 206 may be implemented in software (computer program).

[0179] In the former case, such hardware may be a circuit incorporated into an image encoding or decoding device such as an imaging device, or it may be a circuit incorporated into an image encoding or decoding device that receives images from an external device such as an imaging device or a server device.

[0180] In the latter case, such a computer program may be stored in the memory of an image encoding or decoding device, such as an imaging device, or in memory accessible to an image encoding or decoding device supplied by an external device, such as an imaging device or a server device. A device (computer device) capable of reading and executing such a computer program from memory is applicable to the image encoding device and image decoding device described above. An example of the hardware configuration of such a computer device will be explained using the block diagram in Figure 5.

[0181] The CPU 501 executes various processes using computer programs and data stored in the RAM 502 and ROM 503. In doing so, the CPU 501 controls the operation of the entire computer system and executes or controls the various processes described above as being performed by the image encoding device and image decoding device in each embodiment and modification.

[0182] RAM 502 has an area for storing computer programs and data loaded from the external storage device 506, and an area for storing data acquired from the outside via the I / F (interface) 507. Furthermore, RAM 502 has a work area (such as frame memory) used by the CPU 501 when executing various processes. In this way, RAM 502 can provide various areas as appropriate.

[0183] ROM503 stores configuration data for the computer device, computer programs and data related to the startup of the computer device, computer programs and data related to the basic operation of the computer device, and so on.

[0184] The operation unit 504 is a user interface such as a keyboard, mouse, or touch panel, and allows the user to input various instructions to the CPU 501 through operation.

[0185] The display unit 505 has an LCD screen or a touch panel screen and displays the processing results from the CPU 501 as images, text, etc. The display unit 505 may also be a projection device such as a projector that projects images and text.

[0186] The external storage device 506 is a large-capacity information storage device such as a hard disk drive. The external storage device 506 stores the OS (operating system), computer programs and data for the CPU 501 to execute the various processes described above as being performed by the image encoding device and image decoding device, etc. The external storage device 506 also stores information treated as known information in the above description (encoding tables, tables, etc.). The external storage device 506 may also store the data to be encoded (input images, two-dimensional data arrays, etc.).

[0187] Computer programs and data stored in the external storage device 506 are loaded into the RAM 502 as appropriate according to the control of the CPU 501 and become subject to processing by the CPU 501. The above-mentioned storage unit 103 and frame memories 108 and 206 can be implemented using RAM 502, ROM 503, external storage device 506, etc.

[0188] The I / F507 can be connected to networks such as LANs and the Internet, as well as other devices such as projection and display devices. This computer device can acquire and transmit various types of information via the I / F507.

[0189] The CPU 501, RAM 502, ROM 503, control unit 504, display unit 505, external storage device 506, and I / F 507 are all connected to the system bus 508.

[0190] In the above configuration, when the power to the computer device is turned ON, the CPU 501 executes the boot program stored in the ROM 503, loads the OS stored in the external storage device 506 into the RAM 502, and starts the OS. As a result, the computer device becomes capable of communication via the I / F 507. Then, under the control of the OS, the CPU 501 loads the encoding-related application from the external storage device 506 into the RAM 502 and executes it, so the CPU 501 functions as the various functional units (excluding the storage unit 103 and the frame memory 108) shown in Figures 1 and 13. In other words, the computer device functions as the image encoding device described above. On the other hand, under the control of the OS, the CPU 501 loads the decoding-related application from the external storage device 506 into the RAM 502 and executes it, so the CPU 501 functions as the various functional units (excluding the frame memory 206) shown in Figures 2 and 16. In other words, the computer device functions as the image decoding device described above.

[0191] In this embodiment, we have explained that a computer device having the configuration shown in Figure 5 is applicable to an image encoding device and an image decoding device. However, the hardware configuration of a computer device applicable to an image encoding device and an image decoding device is not limited to the hardware configuration shown in Figure 5. Furthermore, the hardware configuration of a computer device applied to an image encoding device and the hardware configuration of a computer device applied to an image decoding device may be the same or different.

[0192] Furthermore, the numerical values, processing timings, processing order, processing entity, data (information) destination / source / storage location, etc., used in each of the above embodiments and modifications are given as examples for the purpose of providing a concrete explanation, and are not intended to limit the scope to such examples.

[0193] Furthermore, some or all of the embodiments and modified examples described above may be used in appropriate combinations. Alternatively, some or all of the embodiments and modified examples described above may be used selectively.

[0194] (Other embodiments) The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit (e.g., an ASIC) that implements one or more functions.

[0195] The invention is not limited to the embodiments described above, and various modifications and variations are possible without departing from the spirit and scope of the invention. Accordingly, claims are attached to disclose the scope of the invention. [Explanation of Symbols]

[0196] 102: Splitting Unit 104: Prediction Unit 105: Transformation / Quantization Unit 106: Inverse Quantization / Inverse Transformation Unit 107: Image Playback Unit 108: Frame Memory 1309: In-Loop Filter Unit 110: Encoding Unit 111: Integrated Encoding Unit 1313: Determination Unit 150: Control Unit

Claims

1. An encoding means that encodes an image using a prediction mode in block units, A determination means for determining the intensity of a deblocking filter process performed on the boundary between a first block in the image and a second block adjacent to the first block in the image, Processing means that performs the deblocking filter processing on the boundary according to the intensity determined by the determination means. Equipped with, The encoding means has the following prediction mode for the target block: A first prediction mode in which the predicted pixels of the target block are derived using pixels in the image including the target block, A second prediction mode in which the predicted pixels of the target block are derived using pixels in another image different from the image containing the target block, A third prediction mode is provided in which, for a first region in the target block, predicted pixels are derived using pixels in the image containing the target block, without using pixels in another image different from the image containing the target block, and for a second region in the target block different from the first region, predicted pixels are derived using pixels in another image different from the image containing the target block, without using pixels in the image containing the target block. It is possible to use one of several prediction modes, including Each of the first and second regions in the target block can have a shape different from a rectangle. The aforementioned determination means is If at least one of the first block and the second block is a block to which the first prediction mode is applied, a first intensity is determined as the intensity for the deblocking filter process performed on the boundary. If at least one of the first block and the second block is a block to which the third prediction mode is applied, the second intensity is determined as the intensity for the deblocking filtering process performed on the boundary. An image coding device characterized by the following:

2. The target block has four sides, including a first side, a second side perpendicular to the first side, and a third side parallel to the first side and perpendicular to the second side. The second region in the target block can be a region defined by at least a part of the first edge, the second edge, and the third edge. The image coding apparatus according to feature 1.

3. The image encoding device according to Claim 1, characterized in that the first region is a region corresponding to a triangular shape including at least one vertex of the target block, and the second region is a region corresponding to a trapezoidal shape including at least another vertex located diagonally opposite to the aforementioned vertex of the target block.

4. The image coding apparatus according to claim 1, characterized in that the second intensity is the same as the first intensity.

5. The image coding apparatus according to claim 1, characterized in that the second intensity is an intensity between the first intensity and the other intensity.

6. The image coding apparatus according to claim 1, characterized in that, if at least one of the first block and the second block is a block to which the third prediction mode is applied, the second intensity is determined to be the same as the first intensity for predicted pixels obtained from intra-prediction at the boundary, and the second intensity is determined to be less than the first intensity for predicted pixels obtained from inter-prediction at the boundary.

7. The image encoding apparatus according to claim 1, characterized in that the determination means determines the intensity for the luminance component and the intensity for the color difference component.

8. The image coding apparatus according to claim 1, characterized in that the intensity is a bS value.

9. The aforementioned determination means is If at least one of the first block and the second block is a block to which the first prediction mode is applied, the coefficient of the deblocking filter for the first intensity correction intensity is determined. If at least one of the first block and the second block is a block to which the third prediction mode is applied, the coefficient of the deblocking filter for the second intensity correction intensity is determined. The image coding apparatus according to feature 1.

10. The aforementioned determination means is If at least one of the first block and the second block is a block to which the first prediction mode is applied, the filter length of the deblocking filter of the first intensity correction intensity is determined. If at least one of the first block and the second block is a block to which the third prediction mode is applied, the filter length of the deblocking filter for the second intensity correction intensity is determined. The image coding apparatus according to feature 1.

11. A decoding means that decodes an image using a block-based prediction mode, A determination means for determining the intensity of a deblocking filter process performed on the boundary between a first block in the image and a second block adjacent to the first block in the image, Processing means that performs the deblocking filter processing on the boundary according to the intensity determined by the determination means. Equipped with, The decoding means has the following prediction mode for the target block in the image: A first prediction mode in which the predicted pixels of the target block are derived using the pixels in the image including the target block, A second prediction mode in which the predicted pixels of the target block are derived using pixels in another image different from the image containing the target block, A third prediction mode is provided in which, for a first region in the target block, predicted pixels are derived using pixels in the image containing the target block, without using pixels in another image different from the image containing the target block, and for a second region in the target block different from the first region, predicted pixels are derived using pixels in another image different from the image containing the target block, without using pixels in the image containing the target block. It is possible to use one of several prediction modes, including Each of the first and second regions in the target block can have a shape different from a rectangle. The aforementioned determination means is If at least one of the first block and the second block is a block to which the first prediction mode is applied, a first intensity is determined as the intensity for the deblocking filter process performed on the boundary. If at least one of the first block and the second block is a block to which the third prediction mode is applied, the second intensity is determined as the intensity for the deblocking filtering process performed on the boundary. An image decoding device characterized by the following features.

12. The target block has four sides, including a first side, a second side perpendicular to the first side, and a third side parallel to the first side and perpendicular to the second side. The second region in the target block can be a region defined by at least a part of the first edge, the second edge, and the third edge. The image decoding device according to feature 11.

13. The image decoding device according to claim 11, characterized in that the first region is a region corresponding to a triangular shape including at least one vertex of the target block, and the second region is a region corresponding to a trapezoidal shape including at least another vertex located diagonally opposite to the aforementioned vertex of the target block.

14. The image decoding apparatus according to claim 11, characterized in that the second intensity is the same as the first intensity.

15. The image decoding apparatus according to claim 11, characterized in that the second intensity is an intensity between the first intensity and the other intensity.

16. The image decoding apparatus according to claim 11, characterized in that, if at least one of the first block and the second block is a block to which the third prediction mode is applied, the second intensity is determined to be the same as the first intensity for predicted pixels obtained from intra-prediction at the boundary, and the second intensity is determined to be less than the first intensity for predicted pixels obtained from inter-prediction at the boundary.

17. The encoding process involves encoding the image using a prediction mode in block units, A determination step of determining the intensity of a deblocking filter process performed on the boundary between a first block in the image and a second block adjacent to the first block in the image, A processing step in which the deblocking filter process on the boundary is performed according to the intensity determined in the determination step. It has, In the encoding process, the prediction mode for the target block is: A first prediction mode in which the predicted pixels of the target block are derived using pixels in the image including the target block, A second prediction mode in which the predicted pixels of the target block are derived using pixels in another image different from the image containing the target block, A third prediction mode is provided in which, for a first region in the target block, predicted pixels are derived using pixels in the image containing the target block, without using pixels in another image different from the image containing the target block, and for a second region in the target block different from the first region, predicted pixels are derived using pixels in another image different from the image containing the target block, without using pixels in the image containing the target block. It is possible to use one of several prediction modes, including Each of the first and second regions in the target block can have a shape different from a rectangle. In the aforementioned decision-making process, If at least one of the first block and the second block is a block to which the first prediction mode is applied, a first intensity is determined as the intensity for the deblocking filter process performed on the boundary. If at least one of the first block and the second block is a block to which the third prediction mode is applied, the second intensity is determined as the intensity for the deblocking filtering process performed on the boundary. An image encoding method characterized by the following.

18. A decoding process that decodes the image using a block-based prediction mode, A determination step of determining the intensity of a deblocking filter process performed on the boundary between a first block in the image and a second block adjacent to the first block in the image, A processing step in which the deblocking filter process on the boundary is performed according to the intensity determined in the determination step. It has, In the decoding step, the prediction mode for the target block in the image is: A first prediction mode in which the predicted pixels of the target block are derived using the pixels in the image including the target block, A second prediction mode in which the predicted pixels of the target block are derived using pixels in another image different from the image containing the target block, A third prediction mode is provided in which, for a first region in the target block, predicted pixels are derived using pixels in the image containing the target block, without using pixels in another image different from the image containing the target block, and for a second region in the target block different from the first region, predicted pixels are derived using pixels in another image different from the image containing the target block, without using pixels in the image containing the target block. It is possible to use one of several prediction modes, including Each of the first and second regions in the target block can have a shape different from a rectangle. In the aforementioned decision-making process, If at least one of the first block and the second block is a block to which the first prediction mode is applied, a first intensity is determined as the intensity for the deblocking filter process performed on the boundary. If at least one of the first block and the second block is a block to which the third prediction mode is applied, the second intensity is determined as the intensity for the deblocking filtering process performed on the boundary. An image decoding method characterized by the following:

19. A computer program for causing a computer to function as one of the means of an image encoding apparatus according to any one of claims 1 to 10.

20. A computer program for causing a computer to function as one of the means of an image decoding apparatus according to any one of claims 11 to 16.