Image Encoding Device, Image Decoding Device, and Program

By generating an error map to determine optimal orthogonal transformations, the image encoding and decoding devices enhance coding efficiency and reduce information transfer, addressing the inefficiencies in existing technologies.

JP7704929B2Active Publication Date: 2025-07-08NIPPON HOSO KYOKAI
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024074523
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-05-01
Publication Date
2025-07-08
Estimated Expiration
2038-08-15

AI Technical Summary

Technical Problem

Existing image encoding and decoding technologies fail to apply the optimal type of orthogonal transformation based on the energy distribution in the prediction residual, leading to decreased coding efficiency, and methods like KLT increase the amount of information to be transmitted.

Method used

The image encoding and decoding devices utilize an evaluation unit to generate an error map indicating the distribution of errors in the predicted image, allowing for the determination of an optimal orthogonal transformation, and include adaptive transformation generation units to perform orthogonal transformations based on this map, reducing the need for additional information transfer.

Benefits of technology

This approach enables the application of an adaptive orthogonal transformation that efficiently concentrates the energy of the prediction residual, improving encoding efficiency and reducing the amount of information to be transmitted.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007704929000007
    Figure 0007704929000007
  • Figure 0007704929000008
    Figure 0007704929000008
  • Figure 0007704929000009
    Figure 0007704929000009
Patent Text Reader

Abstract

To improve coding efficiency.SOLUTION: An image encoding device (1) encodes a target image in block units obtained by dividing a current image in frame units. The image encoding device (1) includes: a prediction unit (172) for predicting the target image using a plurality of reference images to generate a predicted image; an evaluation unit (180) for generating map information indicating a distribution of an error in the predicted image by evaluating similarity between the plurality of reference images in pixel units; a subtraction unit (110) for calculating a prediction residual indicating a difference between the target block and the prediction image in pixel units; a determination unit (190) for determining an orthogonal transformation to be applied to the prediction residual on the basis of the map information; and a transformation unit (121) for performing an orthogonal transformation process on the prediction residual by the determined orthogonal transformation.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image encoding device, an image decoding device, and a program.

Background Art

[0002] Conventionally, in an image encoding device that encodes a target image in block units obtained by dividing a current image in frame units, a predicted image is generated by predicting the target image using a plurality of reference images, and an orthogonal transformation process is performed on a prediction residual indicating the difference between the target image and the predicted image to calculate a transformation coefficient. A method of quantizing and entropy encoding the transformation coefficient to output encoded data is known.

[0003] Similarly to the image encoding device, the image decoding device generates a predicted image by predicting the target image using a plurality of reference images. The image decoding device decodes the encoded data to obtain a transformation coefficient and performs inverse quantization, performs an inverse orthogonal transformation process on the transformation coefficient after inverse quantization to calculate a prediction residual, and decodes the target image by synthesizing the predicted image and the prediction residual.

[0004] In HEVC, two types of orthogonal transformations applicable to the transformation process (orthogonal transformation process and inverse orthogonal transformation process), namely DCT-2 and DST-7, are defined (see Non-Patent Document 1). Specifically, in HEVC, based on the block size of the target image and the intra prediction mode applied to the target image, it is determined which type of orthogonal transformation to apply among the two types of orthogonal transformations.

Prior Art Documents

Non-Patent Documents

[0005]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0006] However, just determining the type of orthogonal transformation based on the block size of the target image and the mode of intra prediction applied to the target image cannot apply the optimal type of orthogonal transformation according to the energy distribution in the prediction residual. For example, even if DCT-2 is originally more efficient in concentrating the energy of the prediction residual, there is a problem that the coding efficiency decreases because DST-7 may be determined as the orthogonal transformation to be applied.

[0007] Also, KLT is cited as a method of orthogonal transformation that more efficiently concentrates the energy of the prediction residual. However, since the image decoding device side requires information for the inverse processing of KLT performed by the image encoding device, the amount of information to be transmitted increases, resulting in a problem that the coding efficiency decreases.

[0008] Therefore, an object of the present invention is to provide an image encoding device, an image decoding device, and a program that can improve coding efficiency.

Means for Solving the Problems

[0009] The image encoding device according to the first feature is an image encoding device that encodes a target image in block units obtained by dividing a current image in frame units, and includes a prediction unit that generates a prediction image by predicting the target image using a plurality of reference images, an evaluation unit that generates map information indicating the distribution of errors in the prediction image by evaluating the similarity between the plurality of reference images in pixel units, a subtraction unit that calculates a prediction residual indicating the difference between the target block and the prediction image in pixel units, a determination unit that determines an orthogonal transformation to be applied to the prediction residual based on the map information, and a conversion unit that performs an orthogonal transformation process on the prediction residual by the determined orthogonal transformation. The gist is that it comprises these components. The image encoding device according to another feature is an image encoding device that encodes a target image in block units obtained by dividing a current image in frame units, and includes a prediction unit that generates a prediction image by predicting the target image using a plurality of reference images, and an evaluation unit that calculates the sum of absolute differences indicating the similarity between the plurality of reference images in units of regions smaller than the block and consisting of a plurality of pixels, and the gist is that the evaluation unit controls the encoding process based on the sum of absolute differences calculated in the region units.

[0010] The image decoding apparatus according to the second feature is an image decoding apparatus that decodes a target image in block units obtained by dividing a current image in frame units, and includes a decoding unit that acquires transform coefficients by decoding encoded data, a prediction unit that predicts the target image using a plurality of reference images to generate a predicted image, an evaluation unit that generates map information indicating the distribution of errors in the predicted image by evaluating the similarity between the plurality of reference images in pixel units, a determination unit that determines an inverse orthogonal transform to be applied to the transform coefficients based on the map information, and an inverse transform unit that performs an inverse orthogonal transform process on the transform coefficients by the determined inverse orthogonal transform. The gist is that it is provided with these components. The image decoding apparatus according to another feature is an image decoding apparatus that decodes a target image in block units obtained by dividing a current image in frame units, and includes a prediction unit that predicts the target image using a plurality of reference images to generate a predicted image, and an evaluation unit that calculates the sum of absolute differences indicating the similarity between the plurality of reference images in a unit area smaller than the block and composed of a plurality of pixels. The gist is that the evaluation unit controls the decoding process based on the sum of absolute differences calculated in the unit area.

[0011] The program according to the third feature causes a computer to function as the image encoding apparatus according to the first feature.

[0012] The program according to the fourth feature causes a computer to function as the image decoding apparatus according to the second feature.

Advantages of the Invention

[0013] According to the present invention, it is possible to provide an image encoding apparatus, an image decoding apparatus, and a program that can improve the encoding efficiency.

Brief Description of the Drawings

[0014]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Mode for Carrying Out the Invention

[0015] With reference to the drawings, an image encoding device and an image decoding device according to the embodiments will be described. In the following description of the drawings, the same or similar parts are denoted by the same or similar reference numerals.

[0016] <First Embodiment> The image encoding device and the image decoding device according to the first embodiment will be described. The image encoding device and the image decoding device according to the first embodiment perform encoding and decoding of moving images typified by MPEG.

[0017] (Image Encoding Device) FIG. 1 is a diagram showing the configuration of an image encoding device 1 according to the first embodiment. As shown in FIG. 1, the image encoding device 1 includes a block division unit 100, a subtraction unit 110, a conversion / quantization unit 120, an entropy encoding unit 130, an inverse quantization / inverse conversion unit 140, a synthesis unit 150, a memory 160, a prediction unit 170, an evaluation unit 180, and a determination unit 190.

[0018] The block division unit 100 divides an input image in units of frames (or pictures) constituting a moving image into small block-shaped regions, and outputs the blocks obtained by the division to the subtraction unit 110. The size of the blocks is, for example, 32×32 pixels, 16×16 pixels, 8×8 pixels, or 4×4 pixels, etc. The shape of the blocks is not limited to a square and may be a rectangle. The blocks are units for the image encoding device 1 to perform encoding and units for the image decoding device 2 to perform decoding.

[0019] The subtraction unit 110 calculates a prediction residual indicating the difference in pixel units between the block input from the block division unit 100 and the predicted image (predicted block) obtained by the prediction unit 170 predicting the block. Specifically, the subtraction unit 110 calculates the prediction residual by subtracting each pixel value of the predicted image from each pixel value of the block, and outputs the calculated prediction residual to the conversion / quantization unit 120.

[0020] The conversion / quantization unit 120 performs orthogonal transformation processing and quantization processing in block units. The conversion / quantization unit 120 includes a conversion unit 121 and a quantization unit 122.

[0021] The conversion unit 121 performs orthogonal transformation processing on the prediction residual input from the subtraction unit 110 to calculate conversion coefficients, and outputs the calculated conversion coefficients to the quantization unit 122. Orthogonal transformation refers to, for example, discrete cosine transform (DCT), discrete sine transform (DST), Karhunen-Loeve transform (KLT), etc. In the first embodiment, the conversion unit 121 performs orthogonal transformation processing by KLT.

[0022] The quantization unit 122 quantizes the conversion coefficients input from the conversion unit 121 using quantization parameters (Qp) and a quantization matrix, and outputs the quantized conversion coefficients to the entropy encoding unit 130 and the inverse quantization / inverse transformation unit 140. Note that the quantization parameter (Qp) is a parameter that is commonly applied to each conversion coefficient within a block and determines the coarseness of quantization. The quantization matrix is a matrix having quantization values as elements when quantizing each conversion coefficient.

[0023] The entropy encoding unit 130 performs entropy encoding on the conversion coefficients input from the quantization unit 122, performs data compression to generate encoded data (bitstream), and outputs the encoded data to the outside of the image encoding apparatus 1. For entropy encoding, Huffman coding, CABAC (Context-based Adaptive Binary Arithmetic Coding), etc. can be used. Note that the entropy encoding unit 130 receives control information related to prediction input from the prediction unit 170 and also performs entropy encoding of the input control information.

[0024] The inverse quantization / inverse transformation unit 140 performs inverse quantization processing and inverse orthogonal transformation processing in block units. The inverse quantization / inverse transformation unit 140 includes an inverse quantization unit 141 and an inverse transformation unit 142.

[0025] The inverse quantization unit 141 performs an inverse quantization process corresponding to the quantization process performed by the quantization unit 122. Specifically, the inverse quantization unit 141 restores the conversion coefficients by inverse quantizing the conversion coefficients input from the quantization unit 122 using the quantization parameters (Qp) and the quantization matrix, and outputs the restored conversion coefficients to the inverse transformation unit 142.

[0026] The inverse transformation unit 142 performs an inverse orthogonal transformation process corresponding to the orthogonal transformation process performed by the transformation unit 121. For example, when the transformation unit 121 performs a discrete cosine transform, the inverse transformation unit 142 performs an inverse discrete cosine transform. The inverse transformation unit 142 performs an inverse orthogonal transformation process on the transformation coefficients input from the inverse quantization unit 141 to restore the prediction residual, and outputs the restored prediction residual, which is the restored prediction residual, to the synthesis unit 150.

[0027] The synthesis unit 150 synthesizes the restored prediction residual input from the inverse transformation unit 142 and the prediction image input from the prediction unit 170 on a pixel-by-pixel basis. The synthesis unit 150 adds the pixel values of each pixel of the restored prediction residual and the pixel values of each pixel of the prediction image to reconstruct (decode) the block, and outputs the decoded image in block units to the memory 160. Such a decoded image may be referred to as a reconstructed image.

[0028] The memory 160 stores the decoded image input from the synthesis unit 150. The memory 160 stores the decoded image in frame units. The memory 160 outputs the stored decoded image to the prediction unit 170. Note that a loop filter may be provided between the synthesis unit 150 and the memory 160.

[0029] The prediction unit 170 performs prediction in block units. The prediction unit 170 includes an intra prediction unit 171, an inter prediction unit 172, and a switching unit 173.

[0030] The intra prediction unit 171 generates an intra prediction image by referring to the decoded pixel values around the block to be predicted among the decoded images stored in the memory 160, and outputs the generated intra prediction image to the switching unit 173. Also, the intra prediction unit 171 selects the optimal intra prediction mode to be applied to the target block from among a plurality of intra prediction modes, and performs intra prediction using the selected intra prediction mode. The intra prediction unit 171 outputs control information regarding the selected intra prediction mode to the entropy encoding unit 130. Note that the intra prediction modes include Planar prediction, DC prediction, and directional prediction.

[0031] The inter prediction unit 172 uses the decoded image stored in the memory 160 as a reference image to calculate a motion vector by a method such as block matching, predicts a block to be predicted to generate an inter prediction image, and outputs the generated inter prediction image to the switching unit 173. The inter prediction unit 172 selects an optimal inter prediction method from among inter predictions using a plurality of reference images (typically, bidirectional prediction) and inter predictions using one reference image (unidirectional prediction), and performs inter prediction using the selected inter prediction method. The inter prediction unit 172 outputs control information related to inter prediction (such as information on the inter prediction method and motion vector) to the entropy encoding unit 130. When using a plurality of reference images to generate an inter prediction image, the inter prediction unit 172 outputs the plurality of reference images to the evaluation unit 180.

[0032] Note that the prediction performed using a plurality of reference images is typically bidirectional prediction in inter prediction, but is not limited thereto. The prediction unit 170 may perform prediction by intra block copy using a plurality of reference images. In intra block copy, a reference image within the same frame as the current frame is used for predicting a block within the current frame. When performing prediction by intra block copy using a plurality of reference images, the prediction unit 170 outputs the plurality of reference images to the evaluation unit 180.

[0033] The switching unit 173 switches between the intra prediction image input from the intra prediction unit 171 and the inter prediction image input from the inter prediction unit 172, and outputs one of the prediction images to the subtraction unit 110 and the synthesis unit 150.

[0034] The evaluation unit 180 generates map information indicating the distribution of errors in the prediction image generated using the plurality of reference images by evaluating the similarity between the plurality of reference images input from the prediction unit 170 on a pixel-by-pixel basis, and outputs the generated map information to the determination unit 190. Details of the evaluation unit 180 will be described later.

[0035] Based on the map information input from the evaluation unit 180, the determination unit 190 determines an orthogonal transformation to be applied to the prediction residual corresponding to the prediction image whose prediction accuracy has been evaluated by the evaluation unit 180, and outputs the determined orthogonal transformation to the transformation unit 121 and the inverse transformation unit 142. In the first embodiment, the determination unit 190 includes an adaptive transformation generation unit 191 that generates an orthogonal transformation in the vertical direction and an orthogonal transformation in the horizontal direction based on the map information. The transformation unit 121 performs an orthogonal transformation process according to the orthogonal transformation input from the determination unit 190. The inverse transformation unit 142 performs an orthogonal transformation process according to the orthogonal transformation input from the determination unit 190. Details of the adaptive transformation generation unit 191 will be described later.

[0036] (Example of Inter Prediction) FIG. 2 is a diagram showing an example of inter prediction. FIG. 2(a) shows dual prediction as an example of inter prediction, and FIG. 2(b) shows an example of a prediction image generated by dual prediction.

[0037] As shown in FIG. 2(a), dual prediction refers to frames that are temporally before and after the target frame (current frame). In the example of FIG. 2(a), the prediction of a block in the image of the t-th frame is performed by referring to the (t - 1)-th frame and the (t + 1)-th frame. In motion detection, a location (block) similar to the target image block is detected from within the search range set by the system from within the reference frames of the (t - 1)-th and (t + 1)-th frames.

[0038] The detected location is the reference image. Information indicating the relative position of the reference image with respect to the target image block is the arrow shown in the figure and is called a motion vector. The information of the motion vector is encoded by entropy coding together with the frame information of the reference image in the image encoding device 1. On the other hand, the image decoding device detects the reference image based on the information of the motion vector generated by the image encoding device 1.

[0039] As shown in FIGS. 2(a) and 2(b), since the reference images 1 and 2 detected by motion detection are similar partial images aligned within the frame to be referenced with respect to the target image block, they become images similar to the target image block (image to be coded). In the example of FIG. 2(b), the target image block includes a star pattern and a partial circular pattern. The reference image 1 includes a star pattern and an overall circular pattern. The reference image 2 includes a star pattern but does not include a circular pattern.

[0040] A predicted image is generated from such reference images 1 and 2. Note that in general, the prediction process generates a predicted image having the features of each reference image by averaging the reference images 1 and 2 that are partially similar but have different features. However, a more advanced process, for example, generating a predicted image by using in combination a signal enhancement process such as a low-pass filter or a high-pass filter may be used. Here, since the reference image 1 includes a circular pattern and the reference image 2 does not include a circular pattern, when the reference images 1 and 2 are averaged to generate a predicted image, the signal of the circular pattern in the predicted image is halved compared to the reference image 1.

[0041] The difference between the predicted image obtained from the reference images 1 and 2 and the target image block (image to be coded) is the prediction residual. In the prediction residual shown in FIG. 2(b), a large difference occurs only in the shifted portion of the edge of the star pattern and the shifted portion of the circular pattern (hatched portion), but for the other portions, prediction can be performed accurately and the difference becomes small (in the example of FIG. 2(b), no difference occurs).

[0042] The portions where no difference occurs (the non-edge portion of the star pattern and the background portion) are portions where the similarity between the reference image 1 and the reference image 2 is high and accurate prediction has been performed. On the other hand, the portions where a large difference occurs are portions unique to each reference image, that is, portions where the similarity between the reference image 1 and the reference image 2 is extremely low. Therefore, it can be understood that the portions where the similarity between the reference image 1 and the reference image 2 is extremely low have low prediction accuracy and cause a large difference (residual).

[0043] When the prediction residual in which portions with large differences and portions with no differences are mixed is orthogonally transformed in this way and degradation of the transform coefficients due to quantization occurs, such degradation of the transform coefficients propagates throughout the image (block) through inverse quantization and inverse orthogonal transformation. Then, when the prediction residual (restored prediction residual) restored by inverse quantization and inverse orthogonal transformation is combined with the prediction image to reconstruct the target image block, image quality degradation also propagates to portions where high-precision prediction has been performed, such as the non-edge portions and background portions of the star pattern shown in Fig. 2(b).

[0044] (Evaluation Unit) Fig. 3 is a diagram showing an example of the configuration of the evaluation unit 180. As shown in Fig. 3, the evaluation unit 180 includes a difference calculation unit (subtraction unit) 180a, a normalization unit 180b, and an adjustment unit 180c.

[0045] The difference calculation unit 180a calculates the difference value (absolute value of the difference) between the reference image 1 and the reference image 2 in pixel units, and outputs the calculated difference value to the normalization unit 180b. Such a difference value is an example of a value indicating similarity. It can be said that the smaller the difference value, the higher the similarity, and the larger the difference value, the lower the similarity. The difference calculation unit 180a may calculate the difference value after performing a filter process on each reference image. The difference calculation unit 180a may calculate a statistic such as the mean squared error and use such a statistic as the similarity.

[0046] The normalization unit 180b normalizes the difference value input from the difference calculation unit 180a with the maximum difference value within the block (that is, the maximum value of the difference values within the block) and outputs it. The smaller such a difference value, the higher the similarity and the higher the prediction accuracy. On the other hand, the larger the difference value, the lower the similarity and the lower the prediction accuracy (the larger the prediction error).

[0047] The normalization unit 180b normalizes the difference value of each pixel input from the difference calculation unit 180a with the difference value of the pixel having the maximum difference value within the block (that is, the maximum value of the difference values within the block), and outputs the normalized difference value, which is the normalized difference value. Such a normalized difference value can be used as an estimated value representing the magnitude of the prediction error.

[0048] The adjustment unit 180c adjusts the normalized difference value input from the normalization unit 180b based on a quantization parameter (Qp) that determines the coarseness of quantization, and outputs the adjusted normalized difference value. Since the greater the coarseness of quantization, the higher the degree of deterioration of the restored prediction residual, the adjustment unit 180c adjusts the normalized difference value (weight) based on the quantization parameter (Qp).

[0049] The estimated value Rij of the prediction error at each pixel position (ij) output by the evaluation unit 180 can be expressed, for example, by the following equation (1).

[0050] Rij = (abs(Xij - Yij) / maxD × Scale(Qp)) ···(1)

[0051] In equation (1), Xij is the pixel value of pixel ij in the reference image 1, Yij is the pixel value of pixel ij in the reference image 2, and abs is a function for obtaining the absolute value. The difference calculation unit 180a outputs abs(Xij - Yij).

[0052] Also, in equation (1), maxD is the maximum value of the difference value abs(Xij - Yij) within the block. To obtain maxD, it is necessary to obtain the difference values for all pixels within the block, but this process may be omitted and substituted with, for example, the maximum value of adjacent blocks that have already been encoded. Alternatively, maxD may be obtained from the quantization parameter (Qp) using a table that defines the correspondence between the quantization parameter (Qp) and maxD. Alternatively, a fixed value specified in advance in the specification may be used as maxD. The normalization unit 180b outputs abs(Xij - Yij) / maxD.

[0053] In Equation (1), Scale(Qp) is a coefficient that is multiplied according to the quantization parameter (Qp). Scale(Qp) is designed to approach 1.0 when Qp is large and approach 0 when Qp is small, and the degree thereof shall be adjusted by the system. Alternatively, a fixed value specified in advance in the specification may be used as Scale(Qp). Further, for the sake of simplifying the processing, Scale(Qp) may be a fixed value designed according to the system, such as 1.0.

[0054] The adjustment unit 180c outputs abs(Xij - Yij) / maxD × Scale(Qp) as the error estimation value Rij. Further, this Rij may output a weighted value adjusted by a sensitivity function designed according to the system. For example, let abs(Xij - Yij) / maxD × Scale(Qp) = Rij, and let Rij = Clip(Rij, 1.0, 0.0), or Rij = Clip(Rij + offset, 1.0, 0.0) to adjust the sensitivity with an offset. Note that Clip(x, max, min) indicates a process of clipping x to max when x exceeds max and clipping x to min when x is less than min.

[0055] The error estimation value Rij for each pixel position calculated in this way becomes a value within the range from 0 to 1.0. Basically, the error estimation value Rij approaches 1.0 when the difference value of the pixel position ij between the reference images is large (i.e., the prediction accuracy is low), and approaches 0 when the difference value of the pixel position ij between the reference images is small (i.e., the prediction accuracy is high). The evaluation unit 180 outputs two-dimensional map information (hereinafter referred to as "error map") composed of the error estimation values Rij for each pixel position ij within the block.

[0056] (Adaptive Transformation Generation Unit) FIG. 4 is a diagram showing the operation of the adaptive transformation generation unit 191 according to the first embodiment. The adaptive transformation generation unit 191 generates a vertical adaptive orthogonal transformation applied to the prediction residual in the vertical direction and a horizontal adaptive orthogonal transformation applied to the prediction residual in the horizontal direction by principal component analysis using the error map input from the evaluation unit 180.

[0057] Specifically, the adaptive transformation generation unit 191 regards the error map as a set of column vectors to generate a covariance matrix, and calculates the eigenvectors of the generated covariance matrix. The adaptive transformation generation unit 191 outputs the obtained eigenvectors as a vertical adaptive orthogonal transformation. Further, the adaptive transformation generation unit 191 regards the matrix obtained by applying the generated vertical adaptive orthogonal transformation in the vertical direction as a set of row vectors to generate a covariance matrix, and calculates its eigenvectors. The adaptive transformation generation unit 191 outputs the obtained eigenvectors as a horizontal adaptive orthogonal transformation.

[0058] As shown in FIG. 4(a), when the error map has a width w and a height h, the adaptive transformation generation unit 191 regards the error map as w h×1 column vectors, and calculates a covariance matrix Λ h The adaptive transformation generation unit 191 calculates the eigenvectors by diagonalizing the obtained covariance matrix. Here, for the diagonalization of the covariance matrix, for example, the Jacobi method or the like is used to perform iterative calculations. The adaptive transformation generation unit 191 outputs an h×h matrix as a vertical adaptive orthogonal transformation by combining the obtained eigenvectors e0 to e h through transformation.

[0059] Furthermore, as shown in FIG. 4(b), the adaptive transformation generation unit 191 regards the error map as h 1×w row vectors, and calculates a covariance matrix Λ h The adaptive transformation generation unit 191 calculates the eigenvectors by diagonalizing the obtained covariance matrix. The adaptive transformation generation unit 191 outputs a w×w matrix as a horizontal adaptive orthogonal transformation by combining the obtained eigenvectors.

[0060] In video coding methods such as HEVC (see Non-Patent Document 1) and the latest video coding technology (JEM) under consideration by international standardization organizations, for the purpose of performing the conversion process quickly and lightly, it is realized by conversion coefficients with integer precision and bit shifts. Also in the adaptive transformation generation unit 191 according to the present embodiment, the obtained eigenvectors may be approximated to integer coefficients, and the amount of bit shift that does not expand the dynamic range of the conversion coefficients may be specified in advance in the image coding device and the image decoding device.

[0061] As shown in FIG. 1, the conversion unit 121 uses the vertical adaptive orthogonal transformation and the horizontal adaptive orthogonal transformation input from the adaptive transformation generation unit 191 to perform orthogonal transformation processing in the vertical and horizontal directions on the prediction residual generated by the subtraction unit 110, thereby calculating conversion coefficients, and outputs the calculated conversion coefficients to the quantization unit 122.

[0062] Also, the inverse conversion unit 142 uses the vertical adaptive orthogonal transformation and the horizontal adaptive orthogonal transformation input from the adaptive transformation generation unit 191 to perform an inverse orthogonal transformation process corresponding to the orthogonal transformation process performed by the conversion unit 121 on the conversion coefficients input from the inverse quantization unit 141.

[0063] (Image decoding device) FIG. 5 is a diagram showing the configuration of the image decoding device 2 according to the first embodiment. As shown in FIG. 5, the image decoding device 2 includes an entropy code decoding unit 200, an inverse quantization / inverse conversion unit 210, a synthesis unit 220, a memory 230, a prediction unit 240, an evaluation unit 250, and a determination unit 260.

[0064] The entropy code decoding unit 200 decodes the encoded data generated by the image encoding device 1 and outputs the quantized conversion coefficients to the inverse quantization / inverse conversion unit 210. Also, the entropy code decoding unit 200 acquires control information regarding prediction (intra prediction and inter prediction), and outputs the acquired control information to the prediction unit 240.

[0065] The inverse quantization / inverse conversion unit 210 performs inverse quantization processing and inverse orthogonal transformation processing in block units. The inverse quantization / inverse conversion unit 210 includes an inverse quantization unit 211 and an inverse conversion unit 212.

[0066] The inverse quantization unit 211 performs an inverse quantization process corresponding to the quantization process performed by the quantization unit 122 of the image encoding device 1. The inverse quantization unit 211 restores the conversion coefficients by inverse quantizing the quantized conversion coefficients input from the entropy code decoding unit 200 using the quantization parameter (Qp) and the quantization matrix, and outputs the restored conversion coefficients to the inverse conversion unit 212.

[0067] The inverse transform unit 212 performs an inverse orthogonal transform process corresponding to the orthogonal transform process performed by the transform unit 121 of the image encoding apparatus 1. The inverse transform unit 212 performs an inverse orthogonal transform process on the transform coefficients input from the inverse quantization unit 211 to restore the prediction residual, and outputs the restored prediction residual (restored prediction residual) to the synthesis unit 220.

[0068] The synthesis unit 220 reconstructs (decodes) the original block by synthesizing the prediction residual input from the inverse transform unit 212 and the predicted image input from the prediction unit 240 on a pixel-by-pixel basis, and outputs the decoded image in block units to the memory 230.

[0069] The memory 230 stores the decoded image input from the synthesis unit 220. The memory 230 stores the decoded image in frame units. The memory 230 outputs the decoded image in frame units to the outside of the image decoding apparatus 2. Note that a loop filter may be provided between the synthesis unit 220 and the memory 230.

[0070] The prediction unit 240 performs prediction in block units. The prediction unit 240 includes an intra prediction unit 241, an inter prediction unit 242, and a switching unit 243.

[0071] The intra prediction unit 241 refers to the decoded image stored in the memory 230, performs intra prediction according to the control information input from the entropy decoding unit 200, generates an intra prediction image, and outputs the generated intra prediction image to the switching unit 243.

[0072] The inter prediction unit 242 performs inter prediction for predicting a block to be predicted using the decoded image stored in the memory 230 as a reference image. The inter prediction unit 242 performs inter prediction according to the control information (inter prediction method, motion vector information, etc.) input from the entropy decoding unit 200 to generate an inter prediction image, and outputs the generated inter prediction image to the switching unit 243. When the inter prediction unit 242 uses a plurality of reference images to generate an inter prediction image, the inter prediction unit 242 outputs the plurality of reference images to the evaluation unit 250.

[0073] Note that the prediction performed using a plurality of reference images is typically bi-prediction in inter-prediction, but is not limited thereto. The prediction unit 240 may perform prediction by intra-block copy using a plurality of reference images. When performing prediction by intra-block copy using a plurality of reference images, the prediction unit 240 outputs the plurality of reference images to the evaluation unit 250.

[0074] The switching unit 243 switches between the intra-prediction image input from the intra-prediction unit 241 and the inter-prediction image input from the inter-prediction unit 242, and outputs one of the prediction images to the synthesis unit 220.

[0075] The evaluation unit 250 performs the same operation as the evaluation unit 180 (see FIG. 3) of the image encoding device 1. The evaluation unit 250 evaluates the similarity between a plurality of reference images input from the prediction unit 240 in pixel units, generates an error map indicating the distribution of errors in the prediction image generated using the plurality of reference images, and outputs the generated error map to the determination unit 260.

[0076] Based on the error map input from the evaluation unit 250, the determination unit 260 determines the inverse orthogonal transformation to be applied to the prediction residual corresponding to the prediction image whose prediction accuracy has been evaluated by the evaluation unit 250, and outputs the determined inverse orthogonal transformation to the inverse transformation unit 212. The determination unit 260 includes an adaptive transformation generation unit 261 that generates orthogonal transformation in the vertical direction and orthogonal transformation in the horizontal direction based on the error map. The adaptive transformation generation unit 261 performs the same operation (see FIG. 4) as the adaptive transformation generation unit 191 of the image encoding device 1. The inverse transformation unit 212 performs inverse orthogonal transformation processing according to the inverse orthogonal transformation input from the determination unit 260.

[0077] (Summary of the First Embodiment) The image encoding device 1 according to the first embodiment encodes a target image in block units obtained by dividing a current image in frame units. The image encoding device 1 includes a prediction unit 170 that predicts a target image using a plurality of reference images to generate a predicted image, an evaluation unit 180 that generates an error map indicating the distribution of errors in the predicted image by evaluating the similarity between the plurality of reference images in pixel units, a subtraction unit 110 that calculates a prediction residual indicating the difference between a target block and the predicted image in pixel units, a determination unit 190 that determines an orthogonal transformation to be applied to the prediction residual based on the error map, and a transformation unit 121 that performs an orthogonal transformation process on the prediction residual by the determined orthogonal transformation. The determination unit 190 includes an adaptive transformation generation unit 191 that generates an orthogonal transformation based on map information.

[0078] Further, the image decoding device 2 according to the first embodiment decodes a target image in block units obtained by dividing a current image in frame units. The image decoding device 2 includes an entropy decoding unit 200 that obtains transform coefficients by decoding encoded data, a prediction unit 240 that predicts a target image using a plurality of reference images to generate a predicted image, an evaluation unit 250 that generates an error map indicating the distribution of errors in the predicted image by evaluating the similarity between the plurality of reference images in pixel units, a determination unit 260 that determines an inverse orthogonal transformation to be applied to the transform coefficients based on the error map, and an inverse transformation unit 212 that performs an inverse orthogonal transformation process on the transform coefficients by the determined inverse orthogonal transformation. The determination unit 260 includes an adaptive transformation generation unit 261 that generates an inverse orthogonal transformation based on map information.

[0079] As described above, according to the first embodiment, by generating an orthogonal transformation based on an error map indicating the distribution of errors in the predicted image, an optimal orthogonal transformation corresponding to the energy distribution in the prediction residual can be applied.

[0080] In addition, since each of the image encoding device 1 and the image decoding device 2 can generate an orthogonal transformation based on the error map, the image decoding device 2 does not need information for the inverse process of the orthogonal transformation (KLT) performed by the image encoding device 1. Therefore, an increase in the amount of information to be transmitted can be suppressed.

[0081] Therefore, according to the image encoding apparatus 1 and the image decoding apparatus 2 according to the first embodiment, an adaptive orthogonal transformation that efficiently concentrates the energy of the prediction residual can be applied, and the encoding efficiency can be improved.

[0082] <Second Embodiment> The image encoding apparatus 1 and the image decoding apparatus 2 according to the second embodiment will be mainly described with differences from the first embodiment.

[0083] (Image Encoding Apparatus) FIG. 6 is a diagram showing the configuration of the image encoding apparatus 1 according to the second embodiment. As shown in FIG. 6, in the image encoding apparatus 1 according to the second embodiment, the configuration of the determination unit 190 is different from that of the first embodiment. The determination unit 190 includes an adaptive transform generation unit 191 that generates an orthogonal transform by performing principal component analysis of an error map (see FIG. 4), similar to the first embodiment. In the second embodiment, the determination unit 190 further includes a candidate selection unit (first selection unit) 192 and an orthogonal transform selection unit (second selection unit) 193.

[0084] The adaptive transform generation unit 191 outputs an adaptive orthogonal transform (vertical adaptive orthogonal transform and horizontal adaptive orthogonal transform) generated based on the error map to the candidate selection unit 192.

[0085] The candidate selection unit 192 selects one or more orthogonal transform candidates in descending order of the correlation with the adaptive orthogonal transform input from the adaptive transform generation unit 191 from among a plurality of types of predefined orthogonal transforms, and outputs the selected one or more orthogonal transform candidates to the orthogonal transform selection unit 263. The plurality of types of predefined orthogonal transforms are shared by the image encoding apparatus 1 and the image decoding apparatus 2. In the second embodiment, it is assumed that DCT-2, DST-7, DCT-8, DST-1, and DCT-5 are predefined as the plurality of types of orthogonal transforms.

[0086] Specifically, the candidate selection unit 192 evaluates the correlation between the horizontal adaptive orthogonal transform input from the adaptive transform generation unit 191 and each orthogonal transform (DCT-2, DST-7, DCT-8, DST-1, DCT-5). Then, based on the result of the correlation evaluation, the candidate selection unit 192 outputs one or more orthogonal transforms in descending order of correlation among the plurality of types of orthogonal transforms to the orthogonal transform selection unit 193 as adaptive orthogonal transform candidates.

[0087] Note that the number of adaptive orthogonal transform candidates output by the candidate selection unit 192 is predefined to be the same in the image encoding device 1 and the image decoding device 2. Also, the number of adaptive orthogonal transform candidates output by the candidate selection unit 192 may be variable according to the block size of the image block to be encoded, color components (luminance component, chrominance component), etc.

[0088] Furthermore, the candidate selection unit 192 may separately select adaptive orthogonal transform candidates in the horizontal and vertical directions. In such a case, the candidate selection unit 192 further evaluates the correlation between the vertical adaptive orthogonal transform input from the adaptive transform generation unit 191 and each orthogonal transform (DCT-2, DST-7, DCT-8, DST-1, DCT-5), and outputs an adaptive vertical orthogonal transform candidate to the orthogonal transform selection unit 193 in addition to the adaptive horizontal orthogonal transform candidate.

[0089] The orthogonal transform selection unit 193 selects an orthogonal transform to be applied to the prediction residual from among the one or more orthogonal transform candidates input from the candidate selection unit 192, and outputs the selected orthogonal transform to the transform unit 121 and the inverse transform unit 142. Also, the orthogonal transform selection unit 193 outputs an index indicating the type of the selected orthogonal transform to the entropy encoding unit 130.

[0090] For example, the orthogonal transform selection unit 193 calculates the encoding efficiency when each orthogonal transform candidate is applied by simulation, and selects an optimal orthogonal transform according to the result of such simulation. Note that the orthogonal transform selection unit 193 may be configured to select the same type of orthogonal transform in the horizontal and vertical directions, or may be configured to select different types of orthogonal transforms in the horizontal and vertical directions.

[0091] The entropy encoding unit 130 entropy-encodes the adaptive orthogonal transform index input from the orthogonal transform selection unit 193. When different types of orthogonal transforms are selected in the horizontal and vertical directions, the entropy encoding unit 130 entropy-encodes different adaptive orthogonal transform indexes in the horizontal and vertical directions. However, when the adaptive orthogonal transform candidates are composed of one type of orthogonal transform, the orthogonal transform can be uniquely specified in the image decoding apparatus 2, and thus such an index does not have to be entropy-encoded.

[0092] (Image decoding apparatus) FIG. 7 is a diagram showing the configuration of the image decoding apparatus 2 according to the second embodiment. As shown in FIG. 7, the configuration of the decision unit 260 of the image decoding apparatus 2 according to the second embodiment is different from that of the first embodiment. The decision unit 260 includes an adaptive transform generation unit 261 that generates an orthogonal transform by performing principal component analysis of the error map (see FIG. 4), as in the first embodiment. In the second embodiment, the decision unit 260 further includes a candidate selection unit (first selection unit) 262 and an orthogonal transform selection unit (second selection unit) 263.

[0093] The adaptive transform generation unit 261 outputs the adaptive orthogonal transforms (vertical adaptive orthogonal transform and horizontal adaptive orthogonal transform) generated based on the error map to the candidate selection unit 262.

[0094] The candidate selection unit 262 selects one or more orthogonal transform candidates in descending order of correlation with the adaptive orthogonal transform input from the adaptive transform generation unit 261 from among a plurality of types of orthogonal transforms defined in advance, and outputs the selected one or more orthogonal transform candidates to the orthogonal transform selection unit 263. As described above, it is assumed that DCT-2, DST-7, DCT-8, DST-1, and DCT-5 are defined in advance as the plurality of types of orthogonal transforms.

[0095] Specifically, the candidate selection unit 262 evaluates the correlation between the horizontal adaptive orthogonal transform input from the adaptive transform generation unit 261 and each orthogonal transform (DCT-2, DST-7, DCT-8, DST-1, DCT-5). Then, based on the result of the correlation evaluation, the candidate selection unit 262 outputs one or more orthogonal transforms in descending order of correlation among the multiple types of orthogonal transforms as adaptive orthogonal transform candidates to the orthogonal transform selection unit 263.

[0096] As described above, the number of adaptive orthogonal transform candidates output by the candidate selection unit 262 is defined in advance to be the same in the image encoding device 1 and the image decoding device 2. Also, the number of adaptive orthogonal transform candidates output by the candidate selection unit 262 may be variable according to the block size and color components (luminance component, chrominance component) of the image block to be decoded.

[0097] Furthermore, the candidate selection unit 262 may separately select adaptive orthogonal transform candidates in the horizontal and vertical directions. In such a case, the candidate selection unit 262 further evaluates the correlation between the vertical adaptive orthogonal transform input from the adaptive transform generation unit 261 and each orthogonal transform (DCT-2, DST-7, DCT-8, DST-1, DCT-5), and outputs an adaptive vertical orthogonal transform candidate to the orthogonal transform selection unit 263 in addition to the adaptive horizontal orthogonal transform candidate.

[0098] On the other hand, the entropy code decoding unit 200 decodes the index indicating the orthogonal transform selected by the image encoding device 1 from among one or more orthogonal transform candidates, and outputs the index to the orthogonal transform selection unit 263. When selecting different types of orthogonal transforms in the horizontal and vertical directions, the entropy code decoding unit 200 decodes separate adaptive orthogonal transform indices in the horizontal and vertical directions.

[0099] Based on the index input from the entropy decoding unit 200, the orthogonal transformation selection unit 263 selects an inverse orthogonal transformation to be applied to the transform coefficients from among one or more orthogonal transformation candidates input from the candidate selection unit 262, and outputs the selected inverse orthogonal transformation to the inverse transformation unit 212. Note that the orthogonal transformation selection unit 263 may be configured to select the same type of orthogonal transformation in the horizontal and vertical directions, or may be configured to select different types of orthogonal transformations in the horizontal and vertical directions.

[0100] (Summary of the Second Embodiment) In the image encoding apparatus 1 according to the second embodiment, the determination unit 190 includes an adaptive transform generation unit 191 that generates an orthogonal transformation by performing principal component analysis on an error map, a candidate selection unit 192 that selects one or more orthogonal transformation candidates in descending order of correlation with the generated orthogonal transformation from among a plurality of predefined types of orthogonal transformations, and an orthogonal transformation selection unit 193 that selects an orthogonal transformation to be applied to the prediction residual from among the one or more orthogonal transformation candidates. The entropy encoding unit 130 encodes an index indicating the orthogonal transformation selected by the orthogonal transformation selection unit 193 from among the one or more orthogonal transformation candidates.

[0101] Also, in the image decoding apparatus 2 according to the second embodiment, the determination unit 260 includes an adaptive transform generation unit 261 that generates an orthogonal transformation by performing principal component analysis on an error map, a candidate selection unit 262 that selects one or more orthogonal transformation candidates in descending order of correlation with the generated orthogonal transformation from among a plurality of predefined types of orthogonal transformations, and an orthogonal transformation selection unit 263 that selects an inverse orthogonal transformation to be applied to the transform coefficients from among the one or more orthogonal transformation candidates. The entropy decoding unit 200 decodes an index indicating the orthogonal transformation selected by the image encoding apparatus 1 from among the one or more orthogonal transformation candidates. Based on the index, the orthogonal transformation selection unit 263 selects an inverse orthogonal transformation to be applied to the transform coefficients from among the one or more orthogonal transformation candidates.

[0102] Thus, according to the second embodiment, among a plurality of types of orthogonal transforms defined in advance, an orthogonal transform candidate highly correlated with the orthogonal transform (KLT) generated by principal component analysis of the error map is selected, and an orthogonal transform to be applied to the conversion process is selected from among the orthogonal transform candidates. The plurality of types of orthogonal transforms defined in advance (DCT-2, DST-7, DCT-8, DST-1, DCT-5) have a smaller amount of arithmetic processing than KLT.

[0103] Therefore, according to the second embodiment, it is possible to apply an orthogonal transform that efficiently concentrates the energy of the prediction residual, improve the coding efficiency, and reduce the amount of arithmetic processing for the conversion process compared to the first embodiment.

[0104] <The Third Embodiment> Regarding the image encoding device 1 and the image decoding device 2 according to the third embodiment, differences from the first and second embodiments will be mainly described.

[0105] In the second embodiment described above, among a plurality of types of orthogonal transforms defined in advance, an orthogonal transform candidate highly correlated with the orthogonal transform (KLT) generated by principal component analysis of the error map was selected. In contrast, in the third embodiment, an orthogonal transform candidate is selected from among a plurality of types of orthogonal transforms defined in advance by evaluating the feature amount of the error map.

[0106] (Image Encoding Device) FIG. 8 is a diagram showing the configuration of the image encoding device 1 according to the third embodiment. As shown in FIG. 8, the image encoding device 1 according to the third embodiment is different from the second embodiment in that the determination unit 190 includes a feature amount evaluation unit 191a. The feature amount evaluation unit 191a evaluates the feature amount of the error map input from the evaluation unit 180 and outputs the evaluation result to the candidate selection unit 192.

[0107] FIG. 9 is a diagram showing the operation of the feature amount evaluation unit 191a. As shown in FIG. 9, in order to evaluate the energy distribution of the error map, the feature amount evaluation unit 191a divides the error map into four parts in the horizontal direction and four parts in the vertical direction, and for each divided region, the total value Exy (E 00up to E 33 ) is calculated. The feature quantity evaluation unit 191a outputs the evaluated energy distribution E 00 up to E 33 to the candidate selection unit 192.

[0108] The candidate selection unit 192 selects an adaptive orthogonal transformation candidate from a plurality of types of predefined orthogonal transformations (DCT-2, DST-7, DCT-8, DST-1, DCT-5) based on the following conditions, and outputs the selected adaptive orthogonal transformation candidate to the orthogonal transformation selection unit 193.

[0109] (a) For the horizontal direction, the candidate selection unit 192:

[0110] [Number]

[0111] When, DCT-2 and DST-1 are selected as horizontal adaptive orthogonal transformation candidates,

[0112] [Number]

[0113] When, DCT-2 and DST-7 are selected as horizontal adaptive orthogonal transformation candidates,

[0114] [Number]

[0115] When, DCT-2 and DCT-5 are selected as horizontal adaptive orthogonal transformation candidates, If none of them apply, DCT-8 and DST-7 are selected as horizontal adaptive orthogonal transformation candidates.

[0116] (b) For the vertical direction, the candidate selection unit 192:

[0117] [Number]

[0118] When it is, DCT-2 and DST-1 are selected as vertical adaptive orthogonal transform candidates,

[0119] [Number]

[0120] When it is, DCT-2 and DST-7 are selected as vertical adaptive orthogonal transform candidates,

[0121] [Number]

[0122] When it is, DCT-2 and DCT-5 are selected as vertical adaptive orthogonal transform candidates, When neither applies, DCT-8 and DST-7 are selected as vertical adaptive orthogonal transform candidates.

[0123] Here, the number of adaptive orthogonal transform candidates output by the candidate selection unit 192 is two for each of the horizontal and vertical directions. However, the number of adaptive orthogonal transform candidates output by the candidate selection unit 192 may be variable according to the block size of the image block to be coded, color components (luminance component, color difference component), etc.

[0124] Similar to the second embodiment, the orthogonal transform selection unit 193 selects an orthogonal transform to be applied to the prediction residual from among the orthogonal transform candidates input from the candidate selection unit 192, and outputs the selected orthogonal transform to the transform unit 121 and the inverse transform unit 142. Also, the orthogonal transform selection unit 193 outputs an index indicating the type of the selected orthogonal transform to the entropy coding unit 130. The entropy coding unit 130 entropy-codes the adaptive orthogonal transform index input from the orthogonal transform selection unit 193.

[0125] (Image decoding device) FIG. 10 is a diagram showing the configuration of the image decoding apparatus 2 according to the third embodiment. As shown in FIG. 10, the image decoding apparatus 2 according to the third embodiment is different from the second embodiment in that the determination unit 260 includes the feature amount evaluation unit 261a. The feature amount evaluation unit 261a performs the same operation as the feature amount evaluation unit 191a of the image encoding apparatus 1 (see FIG. 9).

[0126] The feature amount evaluation unit 261a divides the error map into four parts in the horizontal direction and four parts in the vertical direction in order to evaluate the energy distribution of the error map, and calculates the total value Exy (E 00 to E 33 ) of the error estimation values for each divided region. The feature amount evaluation unit 261a outputs the evaluated energy distribution E 00 to E 33 to the candidate selection unit 262.

[0127] The candidate selection unit 262 selects an adaptive orthogonal transformation candidate from a plurality of types of predefined orthogonal transformations (DCT-2, DST-7, DCT-8, DST-1, DCT-5) based on the same conditions as the candidate selection unit 192 of the image encoding apparatus 1, and outputs the selected adaptive orthogonal transformation candidate to the orthogonal transformation selection unit 263.

[0128] On the other hand, the entropy decoding unit 200 decodes the index indicating the orthogonal transformation selected by the image encoding apparatus 1 from among the orthogonal transformation candidates, and outputs the index to the orthogonal transformation selection unit 263. The orthogonal transformation selection unit 263 selects the inverse orthogonal transformation to be applied to the transformation coefficients from among the orthogonal transformation candidates input from the candidate selection unit 262 based on the index input from the entropy decoding unit 200, and outputs the selected inverse orthogonal transformation to the inverse transformation unit 212.

[0129] (Summary of the Third Embodiment) In the image encoding apparatus 1 according to the third embodiment, the determination unit 190 includes a feature amount evaluation unit 191a that evaluates the feature amount of the error map, a candidate selection unit 192 that selects one or more orthogonal transformation candidates from a plurality of types of orthogonal transformations defined in advance based on the evaluated feature amount, and an orthogonal transformation selection unit 193 that selects an orthogonal transformation to be applied to the prediction residual from among the one or more orthogonal transformation candidates. The entropy encoding unit 130 encodes an index indicating the orthogonal transformation selected by the orthogonal transformation selection unit 193 from among the one or more orthogonal transformation candidates.

[0130] Also, in the image decoding apparatus 2 according to the third embodiment, the determination unit 260 includes a feature amount evaluation unit 261a that evaluates the feature amount of the error map, a candidate selection unit 262 that selects one or more orthogonal transformation candidates from a plurality of types of orthogonal transformations defined in advance based on the evaluated feature amount, and an orthogonal transformation selection unit 263 that selects an inverse orthogonal transformation to be applied to the transformation coefficients from among the one or more orthogonal transformation candidates. The entropy decoding unit 200 decodes an index indicating the orthogonal transformation selected by the image encoding apparatus 1 from among the one or more orthogonal transformation candidates. The orthogonal transformation selection unit 263 selects an inverse orthogonal transformation to be applied to the transformation coefficients from among the one or more orthogonal transformation candidates based on the index.

[0131] As described above, according to the third embodiment, since the feature amount of the error map can be evaluated by simple arithmetic processing, the amount of arithmetic processing can be reduced compared to the second embodiment in which the orthogonal transformation (KLT) is generated by the principal component analysis of the error map.

[0132] Therefore, according to the third embodiment, it is possible to apply an orthogonal transformation that efficiently concentrates the energy of the prediction residual to improve the encoding efficiency, and to reduce the amount of arithmetic processing for analyzing the error map compared to the second embodiment.

[0133] <Other Embodiments> In the above-described first to third embodiments, an example in which conversion processing is separately performed in the vertical direction and the horizontal direction using one-dimensional orthogonal conversion has been described. However, instead of one-dimensional orthogonal conversion, two-dimensional orthogonal conversion may be used to perform the conversion processing in the vertical direction and the horizontal direction together.

[0134] Further, it may be provided by a program that causes a computer to execute each process performed by the image encoding device 1 and a program that causes a computer to execute each process performed by the image decoding device 2. Further, the program may be recorded on a computer-readable medium. By using a computer-readable medium, it is possible to install the program in a computer. Here, the computer-readable medium on which the program is recorded may be a non-transitory recording medium. The non-transitory recording medium is not particularly limited, and may be, for example, a recording medium such as a CD-ROM or a DVD-ROM.

[0135] Further, the circuits that execute each process performed by the image encoding device 1 may be integrated, and the image encoding device 1 may be configured as a semiconductor integrated circuit (chipset, SoC). Similarly, the circuits that execute each process performed by the image decoding device 2 may be integrated, and the image decoding device 2 may be configured as a semiconductor integrated circuit (chipset, SoC).

[0136] As described above, the embodiments have been described in detail with reference to the drawings. However, the specific configuration is not limited to the above, and various design changes and the like can be made without departing from the gist.

Explanation of Signs

[0137] 1: Image encoding device 2: Image decoding device 100: Block division unit 110: Subtraction unit 120: Conversion / quantization unit 121: Conversion unit 122: Quantization unit 130: Entropy encoding unit 140: Inverse quantization and inverse transformation unit 141: Inverse quantization unit 142: Inverse transformation unit 150: Synthesis unit 160: Memory 170: Prediction unit 171: Intra prediction unit 172: Inter prediction unit 173: Switching unit 180: Evaluation unit 180a: Difference calculation unit 180b: Normalization unit 180c: Adjustment unit 190: Decision unit 191: Adaptive transformation generation unit 191a: Feature evaluation unit 192: Candidate selection unit 193: Orthogonal transformation selection unit 200: Entropy code decoding unit 210: Inverse quantization and inverse transformation unit 211: Inverse quantization unit 212: Inverse transformation unit 220: Synthesis unit 230: Memory 240: Prediction unit 241: Intra prediction unit 242: Inter prediction unit 243: Switching unit 250: Evaluation unit 260: Decision unit 261: Adaptive transformation generation unit 261a: Feature evaluation unit 262: Candidate selection unit 263: Orthogonal transformation selection unit

Claims

1. An image encoding apparatus for encoding a target image in block units obtained by dividing a current image in frame units, comprising: a block division unit that performs block division for dividing the current image into the blocks; a prediction unit that generates a predicted image by predicting the target image by inter prediction including bi-prediction using a plurality of reference images; an evaluation unit that calculates, in units of regions smaller than the block and each consisting of a plurality of pixels, a value indicating the similarity between the plurality of reference images after the size of the block is determined by the block division, only when the prediction unit performs the bi-prediction using a reference image before and a reference image after the target image in terms of time; and an image encoding apparatus, wherein the encoding process is controlled based on the value calculated by the evaluation unit in units of regions.

2. An image decoding apparatus for decoding a target image in block units obtained by dividing a current image in frame units, comprising: a prediction unit that generates a predicted image by predicting the target image by inter prediction including bi-prediction using a plurality of reference images; an evaluation unit that calculates, in units of regions smaller than the block and each consisting of a plurality of pixels, a value indicating the similarity between the plurality of reference images for the block whose size is determined by block division on the encoding side, only when the prediction unit performs the bi-prediction using a reference image before and a reference image after the target image in terms of time; and an image decoding apparatus, wherein the decoding process is controlled based on the value calculated by the evaluation unit in units of regions.

3. A program for causing a computer to function as the image encoding apparatus according to Claim 1.

4. A program for causing a computer to function as the image decoding apparatus according to Claim 2.

Citation Information

Patent Citations

  • Prediction decoder

    JP2000059785A

  • Moving picture encoding apparatus and moving picture decoding apparatus

    JP2003204550A

  • Image processing device and method

    WO2010010942A1

  • Methods and apparatus for transform selection in video encoding and decoding

    WO2010087807A1

  • Moving picture coding device and moving picture decoding device

    WO2011080806A1