Correction device and program
By evaluating and rearranging quantized prediction residuals based on similarity, the image encoding and decoding devices enhance entropy coding efficiency in HEVC's conversion skip mode, addressing inefficiencies in energy distribution.
Patent Information
- Application Number
- JP2025074980
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-07-10
AI Technical Summary
In HEVC, the conversion skip mode, which skips orthogonal transformation processing, leads to decreased coding efficiency due to inefficient entropy coding when energy distribution is not concentrated in the low frequency band.
An image encoding and decoding device that evaluates the similarity between reference images and rearranges quantized prediction residuals based on this similarity to prioritize encoding low-similarity regions, enhancing entropy coding efficiency.
Improves encoding efficiency by preferentially encoding regions with low similarity between reference images, ensuring efficient entropy coding and reducing data redundancy.
Smart Images

Figure 2025105866000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a correction device and a program.
Background Art
[0002] A conventional image encoding apparatus that encodes each encoding target block obtained by block-dividing an image performs motion compensation prediction using a plurality of reference images, generates a predicted image corresponding to the encoding target block, and represents the difference in pixel units between the encoding target block and the predicted image Performs orthogonal transformation on the prediction residual to generate orthogonal transformation coefficients, quantizes the orthogonal transformation coefficients, and performs entropy encoding on the quantized orthogonal transformation coefficients.
[0003] Specifically, entropy encoding includes a process called serialization in which two-dimensionally arranged orthogonal transformation coefficients are read out in a predetermined scan order and converted into a one-dimensional orthogonal transformation coefficient sequence. This serialization refers to a process of encoding in order from the first orthogonal transformation coefficient in the one-dimensional orthogonal transformation coefficient sequence.
[0004] Generally, it is known that energy is concentrated in the low frequency band by orthogonal transformation in image encoding, and the energy after quantization (that is, the value of the quantized orthogonal transformation coefficient) becomes zero in the high frequency band. For this reason, conventionally, the quantized orthogonal transformation coefficients are read out in the scan order from low frequency to high frequency, and an end flag is given to the last significant coefficient (non-zero coefficient), so that only the significant coefficients are efficiently encoded (see, for example, Non-Patent Document 1).
Prior Art Documents
Non-Patent Documents
[0005]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] In HEVC (see Non-Patent Document 1), in addition to the mode that performs orthogonal transformation processing, a conversion skip mode that does not perform orthogonal transformation processing can be applied. When the conversion skip mode is applied, since orthogonal transformation for the prediction residual is not performed, it cannot be expected that energy will concentrate in the low frequency band.
[0007] Therefore, in the conversion skip mode, if entropy coding is performed in the same way as the mode that performs orthogonal transformation, efficient entropy coding cannot be performed, and the coding efficiency will decrease.
[0008] Therefore, an object of the present invention is to provide an image coding device, an image decoding device, and a program that can improve coding efficiency when performing motion compensation prediction using a plurality of reference images.
Means for Solving the Problems
[0009] The image encoding device according to the first feature is an image encoding device that encodes an encoding target block in units of blocks obtained by dividing an input image. In a transform skip mode in which the orthogonal transform process of the encoding target block is skipped, a motion compensation prediction unit that generates a prediction image corresponding to the encoding target block by performing motion compensation prediction using a plurality of reference images, an evaluation unit that evaluates the similarity between the plurality of reference images in units of pixels, a quantization prediction residual obtained by quantizing a prediction residual representing a difference in units of pixels between the encoding target block and the prediction image, or a restored prediction residual obtained by inverse quantizing the quantization prediction residual, a determination unit that calculates a similarity representing a correlation between the similarity evaluated by the evaluation unit and determines whether the similarity is equal to or greater than a threshold value, and when the determination unit determines that the similarity is equal to or greater than the threshold value, an encoding unit that encodes the quantization prediction residual in the order corresponding to the similarity evaluated by the evaluation unit.
[0010] The image decoding device according to the second feature is an image decoding device that decodes a decoding target block in units of blocks from encoded data. In a transform skip mode in which the inverse orthogonal transform process of the decoding target block is skipped, a motion compensation prediction unit that generates a prediction image corresponding to the decoding target block by performing motion compensation prediction using a plurality of reference images, an evaluation unit that evaluates the similarity between the plurality of reference images in units of pixels, a decoding unit that decodes the encoded data and obtains a quantization prediction residual representing a difference in units of pixels between the decoding target block and the prediction image, a similarity representing a correlation between the quantization prediction residual obtained by the decoding unit or a restored prediction residual obtained by inverse quantizing the quantization prediction residual and the similarity evaluated by the evaluation unit, a determination unit that calculates the similarity and determines whether the similarity is equal to or greater than a threshold value, and when the determination unit determines that the similarity is equal to or greater than the threshold value, a rearrangement unit that rearranges and outputs the quantization prediction residual in the order corresponding to the similarity evaluated by the evaluation unit.
[0011] The program according to the third feature aims to make a computer function as an image encoding device according to the first feature.
[0012] The program according to the fourth feature aims to make a computer function as an image decoding device according to the second feature.
Advantages of the Invention
[0013] According to the present invention, it is possible to provide an image encoding device, an image decoding device, and a program capable of improving encoding efficiency when performing motion compensation prediction using a plurality of reference images.
Brief Description of the Drawings
[0014]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Embodiments for Carrying Out the Invention
[0015] Hereinafter, an image encoding apparatus and an image decoding apparatus according to an embodiment of the present invention will be described with reference to the drawings. The image encoding apparatus and the image decoding apparatus according to the embodiment of the present invention perform encoding and decoding of moving images represented by MPEG (Moving Picture Experts Group). In the description of the drawings, the same or similar reference numerals are given to the same or similar components.
[0016] <1. Configuration of Image Encoding Apparatus> FIG. 1 is a diagram showing the configuration of an image encoding apparatus 1A according to the embodiment. As shown in FIG. 1, the image encoding apparatus 1A includes a block division unit 100, a subtraction unit 101, a conversion unit 102a, a quantization unit 102b, a switching unit 102c, an entropy encoding unit 103A, an inverse quantization unit 104a, an inverse conversion unit 104b, a switching unit 104c, a synthesis unit 105, an intra prediction unit 106, a loop filter 107, a frame memory 108, a motion compensation prediction unit 109, a switching unit 110, and an evaluation unit 111.
[0017] The block division unit 100 divides the input image in units of frames (or pictures) into small block-shaped regions, and outputs the image blocks to the subtraction unit 101 (and the motion compensation prediction unit 109). The size of the image block is, for example, 32×32 pixels, 16×16 pixels, 8×8 pixels, or 4×4 pixels, etc. However, the shape of the image block is not limited to a square, and may be a rectangle, trapezoid, triangle, L-shape, etc. The image block is a unit for the image encoding device 1A to perform encoding and a unit for the image decoding device 2A (see FIG. 2) to perform decoding, and such an image block is called an encoding target block. Such an image block is also referred to as a coding unit (CU) or a coding block (CB).
[0018] The subtraction unit 101 calculates a prediction residual representing the pixel-by-pixel difference between the encoding target block input from the block division unit 100 and the prediction image (prediction image block) corresponding to the encoding target block. Specifically, the subtraction unit 101 calculates the prediction residual by subtracting each pixel value of the prediction image from each pixel value of the encoding target block, and outputs the calculated prediction residual to the conversion unit 102a. The prediction image is input to the subtraction unit 101 from the intra prediction unit 106 or the motion compensation prediction unit 109 described later via the switching unit 110.
[0019] The conversion unit 102a, the quantization unit 102b, and the switching unit 103c constitute a conversion and quantization unit 102 that performs orthogonal conversion processing and quantization processing in block units.
[0020] The conversion unit 102a calculates orthogonal transformation coefficients for each frequency component by performing an orthogonal transformation on the prediction residual input from the subtraction unit 101, and outputs the calculated orthogonal transformation coefficients to the quantization unit 102b. For the orthogonal transformation, for example, a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loeve transform (KLT), etc. may be used. By such an orthogonal transformation, the residual signal in the pixel region is transformed into the frequency region. In general, the energy of the residual signal in the pixel region is concentrated on the low-frequency side by the orthogonal transformation.
[0021] Note that if irrational numbers or floating-point numbers are used in the operation of the orthogonal transformation, there may be a slight difference in the operation results between devices depending on the operation method and accuracy of the encoding device and the decoding device. To avoid this difference, a transformation that simulates the orthogonal transformation basis with integers or rational numbers may be used. Such a transformation may not have orthogonality, but in this embodiment, the term orthogonal transformation is also used for a transformation that does not have orthogonality, and for the inverse transformation, the term inverse orthogonal transformation is also used for an inverse transformation that does not have orthogonality.
[0022] The switching unit 102c switches the input of the quantization unit 102b between the output of the conversion unit 102a and the output of the subtraction unit 101. Normally, the switching unit 102c selects the orthogonal transformation coefficients, which are the output of the conversion unit 102a, and selects the prediction residual, which is the output of the subtraction unit 101, in the case of a transform skip encoding target block described later. Therefore, in the case of a transform skip encoding target block, the prediction residual and the orthogonal transformation coefficients can be regarded as the same signal. Also, in the case of a transform skip encoding target block, the quantized prediction residual and the quantized orthogonal transformation coefficients can be regarded as the same signal.
[0023] The quantization unit 102b quantizes the orthogonal transform coefficients input from the conversion unit 102a via the switching unit 102c using a quantization parameter (Qp) and a quantization matrix, and generates quantized orthogonal transform coefficients (quantized orthogonal transform coefficients). The quantization parameter (Qp) is a parameter that is commonly applied to each orthogonal transform coefficient within a block, and is a parameter that determines the coarseness of quantization. The quantization matrix is a matrix having quantization values as elements when quantizing each orthogonal transform coefficient. The quantization unit 102b outputs the generated quantized orthogonal transform coefficients to the entropy encoding unit 103A and the inverse quantization unit 104a.
[0024] The entropy encoding unit 103A performs entropy encoding on the quantized orthogonal transform coefficients input from the quantization unit 102b, generates encoded data (bit stream) by data compression of such entropy encoding, and outputs the encoded data to the outside of the image encoding apparatus 1A. For entropy encoding, Huffman coding, CABAC (Context-based Adaptive Binary Arithmetic Coding), or the like can be used.
[0025] Entropy encoding includes a serialization process of reading out the two-dimensionally arrayed quantized orthogonal transform coefficients in a predetermined scan order and converting them into a one-dimensional sequence of quantized orthogonal transform coefficients, and encoding is performed in order from the quantized orthogonal transform coefficient at the head of the one-dimensional sequence of quantized orthogonal transform coefficients. Since energy is concentrated in the low-frequency region by orthogonal transform and the energy (the value of the quantized orthogonal transform coefficient) in the high-frequency region converges to zero by quantization, the quantized orthogonal transform coefficients are read out in the scan order from low frequency to high frequency, and an end flag is given to the last significant coefficient in the scan order, thereby efficiently encoding only the significant coefficients.
[0026] Note that the entropy encoding unit 103A receives information related to prediction from the intra prediction unit 106 and the motion compensation prediction unit 109, and receives information related to filtering processing from the loop filter 107. The entropy encoding unit 103A also performs entropy encoding of this information. Further, when the transform skip mode is applied to the block to be encoded, the entropy encoding unit 103A includes a transform skip flag indicating that the transform skip mode is applied to the block to be encoded in the encoded data.
[0027] The inverse quantization unit 104a, the inverse transform unit 104b, and the switching unit 104c constitute an inverse quantization and inverse transform unit 104 that performs inverse quantization processing and inverse orthogonal transform processing in block units.
[0028] The inverse quantization unit 104a performs inverse quantization processing corresponding to the quantization processing performed by the quantization unit 102b. Specifically, the inverse quantization unit 104a restores the orthogonal transform coefficients by inverse quantizing the quantized orthogonal transform coefficients input from the quantization unit 102b using the quantization parameter (Qp) and the quantization matrix, and outputs the restored orthogonal transform coefficients (restored orthogonal transform coefficients) to the inverse transform unit 104b.
[0029] The switching unit 104c switches the output of the inverse quantization and inverse transform unit 104 between the output of the inverse transform unit 104b and the output of the inverse quantization unit 104a. Normally, the restored orthogonal transform coefficients, which are the output of the inverse transform unit 104b, are selected, and in the case of a transform skip encoding target block described later, the restored prediction residue, which is the output of the inverse quantization unit 104a, is selected. Therefore, in the case of a transform skip encoding target block, the restored prediction residue and the restored orthogonal transform coefficients can be regarded as the same signal. Also, in the case of a transform skip encoding target block, the quantized orthogonal transform coefficients and the quantized prediction error can be regarded as the same signal.
[0030] The inverse transform unit 104b performs an inverse orthogonal transform process corresponding to the orthogonal transform process performed by the transform unit 102a. For example, when the transform unit 102a performs a discrete cosine transform, the inverse transform unit 104b performs an inverse discrete cosine transform. The inverse transform unit 104b performs an inverse orthogonal transform on the restored orthogonal transform coefficients input from the inverse quantization unit 104a to restore the prediction residual, and outputs the restored prediction residual, which is the restored prediction residual, to the synthesis unit 105.
[0031] The synthesis unit 105 synthesizes the restored prediction residual input from the inverse transform unit 104b via the switching unit 104c and the prediction image input from the switching unit 110 on a pixel-by-pixel basis. The synthesis unit 105 adds each pixel value of the restored prediction residual and each pixel value of the prediction image to reconstruct the encoding target block, and outputs the reconstructed block, which is the reconstructed encoding target block, to the intra prediction unit 106 and the loop filter 107.
[0032] The intra prediction unit 106 performs intra prediction using the reconstructed block input from the synthesis unit 105 to generate an intra prediction image, and outputs the intra prediction image to the switching unit 110. Further, the intra prediction unit 106 outputs information such as the selected intra prediction mode to the entropy encoding unit 103A.
[0033] The loop filter 107 performs filter processing as post-processing on the reconstructed block input from the synthesis unit 105, and outputs the reconstructed block after the filter processing to the frame memory 108. Further, the loop filter 107 outputs information regarding the filter processing to the entropy encoding unit 103A. The filter processing includes deblocking filter processing and sample adaptive offset processing.
[0034] The frame memory 108 stores the reconstructed block input from the loop filter 107 as a decoded image in frame units.
[0035] The motion compensation prediction unit 109 performs inter prediction using one or more reconstructed blocks (decoded images) stored in the frame memory 108 as reference images. Specifically, the motion compensation prediction unit 109 calculates a motion vector by a method such as block matching, generates a motion compensation prediction image based on the motion vector, and outputs the motion compensation prediction image to the switching unit 110. Further, the motion compensation prediction unit 109 outputs information regarding the motion vector to the entropy encoding unit 103A.
[0036] The switching unit 110 switches between the intra prediction image input from the intra prediction unit 106 and the motion compensation prediction image input from the motion compensation prediction unit 109, and outputs the prediction image to the subtraction unit 101 and the synthesis unit 105.
[0037] On the other hand, in the transform skip mode in which the orthogonal transform process of the block to be encoded is skipped, the block to be encoded output by the block partitioning unit 100 is quantized without being orthogonally transformed after being converted into a prediction residual in the subtraction unit 101. Specifically, the switching unit 102c selects the output of the subtraction unit 101, and the prediction residual output by the subtraction unit 101 skips the orthogonal transform in the transform unit 102a and is input to the quantization unit 102b. The quantization unit 102b quantizes the prediction residual for the block to be encoded for which the orthogonal transform is skipped (hereinafter referred to as "transform skip block to be encoded"), and outputs the quantized prediction residual, i.e., the quantized prediction residual, to the entropy encoding unit 103A and the inverse quantization unit 104a.
[0038] For the transform skip block to be encoded, the entropy encoding unit 103A performs entropy encoding on the quantized prediction residual input from the quantization unit 102b, performs data compression to generate encoded data (bitstream), and outputs the encoded data to the outside of the image encoding apparatus 1A.
[0039] For the transform skip coding target block, the inverse quantization unit 104a performs an inverse quantization process corresponding to the quantization process performed by the quantization unit 102b. For the transform skip block, by the switching unit 104c selecting the output of the inverse quantization unit 104a, the inverse transform unit 104b skips the inverse orthogonal transform process. Therefore, the signal restored by the inverse quantization unit 104a is input to the synthesis unit 105 as a restored prediction residual without undergoing the inverse orthogonal transform process.
[0040] When the motion compensation prediction unit 109 performs motion compensation prediction using a plurality of reference images for the transform skip coding target block, the evaluation unit 111 evaluates the similarity between the plurality of reference images in terms of pixels, and outputs map information representing the evaluation result to the entropy coding unit 103A. The entropy coding unit 103A encodes the restored prediction residual input from the quantization unit 102b after rearranging it, based on the map information input from the evaluation unit 111, for the transform skip coding target block. Details of the evaluation unit 111 and the entropy coding unit 103A will be described later.
[0041] <2. Configuration of Image Decoding Device> FIG. 2 is a diagram showing the configuration of the image decoding device 2A according to the embodiment. As shown in FIG. 2, the image decoding device 2A includes an entropy decoding unit 200A, an inverse quantization unit 201a, an inverse transform unit 201b, a switching unit 201c, a synthesis unit 202, an intra prediction unit 203, a loop filter 204, a frame memory 205, a motion compensation prediction unit 206, a switching unit 207, and an evaluation unit 208.
[0042] The entropy decoding unit 200A decodes the encoded data generated by the image encoding device 1, and outputs the quantized orthogonal transform coefficients to the inverse quantization unit 201a. Also, the entropy decoding unit 200A decodes the encoded data, acquires information related to prediction (intra prediction and motion compensation prediction) and information related to filter processing, outputs the information related to prediction to the intra prediction unit 203 and the motion compensation prediction unit 206, and outputs the information related to filter processing to the loop filter 204.
[0043] The inverse quantization unit 201a, the inverse transform unit 201b, and the switching unit 201c constitute an inverse quantization and inverse transform unit 201 that performs inverse quantization processing and inverse orthogonal transform processing in block units.
[0044] The inverse quantization unit 201a performs an inverse quantization process corresponding to the quantization process performed by the quantization unit 102b of the image encoding device 1A. The inverse quantization unit 201a inverse quantizes the quantized orthogonal transform coefficients input from the entropy decoding unit 200A using the quantization parameter (Qp) and the quantization matrix to restore the orthogonal transform coefficients, and outputs the restored orthogonal transform coefficients (restored orthogonal transform coefficients) to the inverse transform unit 201b.
[0045] The switching unit 201c switches the output of the inverse quantization and inverse transform unit 201 between the output of the inverse transform unit 201b and the output of the inverse quantization unit 201a. Normally, the output of the inverse transform unit 201b is selected, and in the case of a transform skip decoding target block described later, the output of the inverse quantization unit 201a is selected. Thereby, in the case of a transform skip decoding target block, the restored prediction residual and the restored orthogonal transform coefficients can be regarded as the same signal, and in the case of a transform skip decoding target block, the quantized orthogonal transform coefficients and the quantized prediction residual can be regarded as the same signal.
[0046] The inverse transform unit 201b performs an inverse orthogonal transform process corresponding to the orthogonal transform process performed by the transform unit 102a of the image encoding device 1A. The inverse transform unit 201b performs an inverse orthogonal transform on the restored orthogonal transform coefficients input from the inverse quantization unit 201a to restore the prediction residual, and outputs the restored prediction residual, which is the restored prediction residual, to the synthesis unit 202.
[0047] The synthesis unit 202 synthesizes the restored prediction residual input from the inverse transform unit 201b via the switching unit 201c and the predicted image input from the switching unit 207 in pixel units to reconstruct the original target block, and outputs the reconstructed block to the intra prediction unit 203 and the loop filter 204.
[0048] The intra prediction unit 203 refers to the reconstructed block input from the synthesis unit 202, performs intra prediction based on the intra prediction information input from the entropy decoding unit 200A to generate an intra prediction image, and outputs the intra prediction image to the switching unit 207.
[0049] Based on the filter processing information input from the entropy decoding unit 200A, the loop filter 204 performs the same filter processing as the filter processing performed by the loop filter 107 of the image encoding device 1A on the reconstructed block input from the synthesis unit 202, and outputs the reconstructed block after the filter processing to the frame memory 205.
[0050] The frame memory 205 stores the reconstructed block input from the loop filter 204 in frame units and outputs it as a decoded image to the outside of the image decoding device 2A.
[0051] Using the decoded image stored in the frame memory 205 as a reference image, the motion compensation prediction unit 206 performs motion compensation prediction (inter prediction) based on the motion vector information input from the entropy decoding unit 200A to generate a motion compensation prediction image, and outputs the motion compensation prediction image to the switching unit 207.
[0052] The switching unit 207 switches between the intra prediction image input from the intra prediction unit 203 and the motion compensation prediction image input from the motion compensation prediction unit 206, and outputs the prediction image to the synthesis unit 202.
[0053] On the other hand, for the decoding target block (transform skip decoding target block) to which the transform skip mode is applied, the entropy decoding unit 200A decodes the encoded data generated by the encoding device 1 and outputs the quantized prediction residual (quantized prediction residual) to the inverse quantization unit 201a.
[0054] Regarding the transform skip block, the inverse quantization unit 201a performs an inverse quantization process corresponding to the quantization process performed by the quantization unit 102b of the image encoding apparatus 1A. Regarding the transform skip block, the switching unit 201c selects the output of the inverse quantization unit 201a, so that the inverse transform unit 201b skips the inverse orthogonal transform process. Therefore, the prediction residual (restored prediction residual) restored by the inverse quantization unit 201a is input to the synthesis unit 202 without going through the inverse orthogonal transform process.
[0055] The evaluation unit 208 performs the same operation as the evaluation unit 111 of the image encoding apparatus 1A. Specifically, for the transform skip decoding target block, when the motion compensation prediction unit 206 performs motion compensation prediction using a plurality of reference images, the evaluation unit 208 evaluates the similarity between the plurality of reference images in units of pixels, and outputs map information representing the evaluation result to the entropy decoding unit 200A. The entropy decoding unit 200A decodes the encoded data for the transform skip encoding target block to obtain the quantized prediction residual, and rearranges the quantized prediction residual in the original order based on the map information input from the evaluation unit 208 and outputs it. Details of the evaluation unit 208 and the entropy decoding unit 200A will be described later.
[0056] <3. Motion Compensation Prediction> FIG. 3 is a diagram showing an example of motion compensation prediction. FIG. 4 is a diagram showing an example of a predicted image and a prediction residual. As a simple example of motion compensation prediction, the case of using dual prediction, particularly forward and backward prediction (bidirectional prediction) used in HEVC, will be described.
[0057] As shown in FIG. 3, motion compensation prediction refers to the frames temporally before and after the encoding target frame (current frame). In the example of FIG. 3, motion compensation prediction of a block in the image of the t-th frame, which is the current frame, is performed with reference to the (t - 1)-th frame and the (t + 1)-th frame. Motion compensation detects a location (block) similar to the encoding target block from within the search range set by the system from within the reference frames of the (t - 1) and (t + 1) frames.
[0058] The detected location is a reference image. Information representing the relative position of the reference image with respect to the block to be encoded is an arrow shown in the figure and is called a motion vector. The information of the motion vector is encoded by entropy encoding together with the frame information of the reference image in the image encoding device 1A. On the other hand, the image decoding device 2A detects the reference image based on the information of the motion vector generated by the image encoding device 1A.
[0059] As shown in FIGS. 3 and 4, since the reference images 1 and 2 detected by motion compensation are similar partial images aligned within the frame to be referred to with respect to the block to be encoded, they become images similar to the block to be encoded. In the example of FIG. 4, the block to be encoded includes a star pattern and a partial circular pattern. The reference image 1 includes a star pattern and an overall circular pattern. The reference image 2 includes a star pattern but does not include a circular pattern.
[0060] A predicted image is generated from such reference images 1 and 2. Since the prediction process is a process with a high processing load, it is common to generate a predicted image by averaging the reference images 1 and 2. However, a more advanced process, for example, a signal enhancement process using a low-pass filter or a high-pass filter, may be used in combination to generate a predicted image. Here, since the reference image 1 includes a circular pattern and the reference image 2 does not include a circular pattern, when the reference images 1 and 2 are averaged to generate a predicted image, the signal of the circular pattern in the predicted image is halved compared to the reference image 1.
[0061] The difference between the predicted image obtained from the reference images 1 and 2 and the block to be encoded is the prediction residual. The difference between the predicted image obtained from the reference images 1 and 2 and the target image block (the block to be encoded) is the prediction residual. In the prediction residual shown in FIG. 4, a large difference occurs in the shifted part (hatched part) of the circular pattern, but for the other parts, prediction can be performed accurately and no difference occurs.
[0062] The part of the star pattern where there is no difference is the part with a high similarity between reference image 1 and reference image 2, and is the part where high-precision prediction has been performed. On the other hand, the part where a large difference occurs is the part unique to each reference image, that is, the part with a low similarity between reference image 1 and reference image 2. Therefore, the part with a low similarity between reference image 1 and reference image 2, that is, "region A" in the prediction residual, has low prediction accuracy and can be said to cause a large prediction residual.
[0063] Therefore, in the conversion skip mode, the image encoding device 1A preferentially encodes the part with a low similarity between reference image 1 and reference image 2 as the part with a large prediction residual. Specifically, the entropy encoding unit 103A preferentially encodes significant coefficients by encoding the quantized prediction residual in order from the pixel positions with a low similarity between reference images, so that an end flag can be assigned earlier.
[0064] On the other hand, "region B" in the prediction residual is a part with a high similarity between reference image 1 and reference image 2, but a large prediction residual is generated. For this reason, if only region A, which is the part with a low similarity between reference image 1 and reference image 2, is regarded as the part with a large quantized prediction residual and preferentially encoded, the encoding of "region B" will be postponed and the assignment of the end flag will be delayed.
[0065] In particular, in the encoding of a conventional conversion skip encoding target block used in HEVC, encoding is performed by diagonal scan (a scan method performed in the upper right diagonal direction, similar to the zigzag scan such as MPEG-2). For this reason, when using the conventional diagonal scan, region B is encoded earlier and the end flag is also assigned earlier. That is, in such a case, the end flag is assigned earlier by encoding in the order according to the conventional diagonal scan than by preferentially encoding the part with a low similarity between reference image 1 and reference image 2 as the part with a large prediction residual.
[0066] Therefore, in the embodiment, in the conversion skip mode, the correlation between the map information representing the distribution of the similarity between the reference image 1 and the reference image 2 in terms of pixels and the distribution of the quantized prediction residual, which is the "quantized prediction residual", is obtained. When the correlation is low, encoding is performed according to the conventional scan order, and when the correlation is high, encoding is performed in order from the portion with low similarity based on the map information. Thereby, efficient entropy encoding can be performed, and the encoding efficiency can be improved.
[0067] Hereinafter, an example of obtaining the correlation between the map information representing the distribution of the similarity between the reference image 1 and the reference image 2 in terms of pixels and the distribution of the quantized prediction residual will be described. However, when obtaining the correlation, the restored prediction residual may be used instead of the quantized prediction residual. That is, the correlation between the map information representing the distribution of the similarity between the reference image 1 and the reference image 2 in terms of pixels and the distribution of the restored prediction residual may be obtained.
[0068] <4. Evaluation Unit> FIG. 5 is a diagram showing an example of the configuration of the evaluation unit 111 in the image encoding apparatus 1A. As shown in FIG. 5, the evaluation unit 111 includes a difference calculation unit 111a and a normalization unit 111b.
[0069] The difference calculation unit 111a calculates the difference value (specifically, the absolute value of the difference value; the same applies hereinafter) between the reference images 1 and 2 input from the motion compensation prediction unit 109 in terms of pixels, and outputs the calculated difference value to the normalization unit 111b.
[0070] Here, it can be said that the larger the difference value, the lower the similarity and the lower the prediction accuracy. The smaller the difference value, the higher the similarity and the higher the prediction accuracy. The difference calculation unit 111a may calculate the difference value after performing filter processing on each reference image. The difference calculation unit 111a may calculate a statistic such as the mean squared error and output such a statistic. Calculate a statistic such as the mean squared error and output such a statistic.
[0071] The normalization unit 111b normalizes the difference value in pixel units input from the difference calculation unit 111a with the maximum difference value in the block to be encoded (i.e., the maximum value of the difference values in the block to be encoded) and outputs it. The normalization unit 111b may adjust and output the difference value input from the normalization unit 111b based on at least one of a quantization parameter (Qp) that determines the coarseness of quantization and a quantization matrix to which different quantization values are applied for each orthogonal transformation coefficient.
[0072] The evaluation unit 111 calculates the accuracy Rij in pixel units within the block to be encoded, for example, as shown in the following formula (1).
[0073] Rij = 1 - (abs(Xij - Yij) / maxD × Scale(Qp)) ···(1) However, Xij is the pixel value at the pixel position ij of the reference image 1, Yij is the pixel value at the pixel position ij of the reference image 2, and abs is a function that obtains the absolute value.
[0074] Also, in formula (1), maxD is the maximum value of the difference value abs(Xij - Yij) within the block to be encoded. Although it is necessary to obtain the difference value for all pixel positions within the block to be encoded to obtain maxD, this process may be omitted and substituted with the maximum value of the already encoded adjacent blocks. Alternatively, maxD may be obtained from the quantization parameter (Qp) or the quantization value of the quantization matrix using a table that defines the correspondence relationship between the quantization parameter (Qp) or the quantization value of the quantization matrix and maxD. Alternatively, a fixed value specified in advance in the specification may be used as maxD.
[0075] Also, in formula (1), Scale(Qp) is a coefficient that is multiplied according to the quantization parameter (Qp) or the quantization value of the quantization matrix. Scale(Qp) is designed to approach 1.0 when Qp or the quantization value of the quantization matrix is large and approach 0 when small, and the degree is adjusted by the system. Alternatively, a fixed value specified in advance in the specification may be used as Scale(Qp).
[0076] In addition, when using substitute values such as fixed values for maxD or Scale(Qp), there may be cases where the accuracy Rij exceeds 1.0 or is less than 0. In such cases, clipping may be performed to 1.0 or 0.
[0077] The evaluation unit 111 outputs map information composed of the accuracy Rij at each pixel position ij within the block to be encoded to the entropy encoding unit 103A.
[0078] However, the evaluation unit 111 performs evaluation (calculation of accuracy Rij) only when applying motion compensation prediction using a plurality of reference images to the conversion skip encoding target block, and in other modes, such as unidirectional prediction or intra prediction processing, evaluation (calculation of accuracy Rij) may not be performed.
[0079] Note that the evaluation unit 208 in the image decoding device 2A is configured in the same manner as the evaluation unit 111 in the image encoding device 1A. Specifically, the evaluation unit 208 in the image decoding device 2A includes a difference calculation unit 208a and a normalization unit 208b. The evaluation unit 208 in the image decoding device 2A outputs map information composed of the accuracy Rij at each pixel position ij within the block to be encoded to the entropy decoding unit 200A.
[0080] <5. Entropy Encoding Unit> FIG. 6 is a diagram showing an example of the configuration of the entropy encoding unit 103A. FIG. 7 is a diagram showing an example of the generation of the accuracy index. FIG. 8 is a diagram showing an example of the rearrangement of the quantized prediction residuals.
[0081] As shown in FIG. 6, the entropy encoding unit 103A includes a sorting unit 103a, a prediction residual rearrangement unit 103b, a serialization unit 103c, an accuracy rearrangement unit 103d, a determination unit 103e, a switching unit 103f, and an encoding unit 103g.
[0082] The sorting unit 103a sorts the accuracies Rij in the map information input from the evaluation unit 111 in ascending order. Specifically, as shown in FIG. 7(A), since the accuracies Rij are two-dimensionally arranged in the map information, the sorting unit 103a serializes the map information by, for example, zigzag scanning (scanning order from the upper left to the lower right) to obtain an accuracy column. Then, as shown in FIG. 7(B), the sorting unit 103a sorts the accuracies Rij in ascending order, and outputs accuracy index information in which the index i is associated with the index i and the pixel position (the position of the X coordinate and the position of the Y coordinate) to the prediction residual rearrangement unit 103b and the accuracy rearrangement unit 103d using the accuracy Rij as the index i.
[0083] In the example of FIG. 7(A), the accuracy at the pixel position (x, y) = (2, 2) is the lowest, the accuracy at the pixel position (x, y) = (2, 3) is the second lowest, the accuracy at the pixel position (x, y) = (3, 2) is the third lowest, and the accuracy at the pixel position (x, y) = (3, 3) is the fourth lowest. It can be estimated that the region composed of such pixel positions has low prediction accuracy and generates a large prediction residual. On the other hand, for pixel positions with an accuracy of 1, it can be estimated that the prediction accuracy is high and no prediction residual is generated.
[0084] The prediction residual rearrangement unit 103b rearranges the quantized prediction residuals input from the quantization unit 102b for the transform skip blocks based on the accuracy index information input from the sorting unit 103a. Specifically, the prediction residual rearrangement unit 103b rearranges the quantized prediction residuals on a pixel-by-pixel basis so that the quantized prediction residuals are encoded in order from pixel positions with low accuracies (i.e., pixel positions with low similarity between reference images).
[0085] The quantization prediction residuals shown in FIG. 8(A) are arranged in two dimensions, and the prediction residual rearrangement unit 103b rearranges, as shown in FIG. 8(B), the quantization prediction residuals at pixel positions with low probabilities so as to concentrate them in the upper left region based on the probability index information input from the sorting unit 103a. Here, zigzag scan is assumed as the scan order, and the prediction residual rearrangement unit 103b scans in the order from the upper left region to the lower right region, preferentially serializes the quantization prediction residuals at pixel positions with low probabilities, and outputs a quantization prediction residual sequence in which the quantization prediction residuals are arranged in ascending order of low probability (i.e., low similarity between reference images) to the switching unit 103f and the determination unit 103e. However, the scan order is not limited to the zigzag scan, and a horizontal scan or a vertical scan may be used. The prediction residual rearrangement unit 103b may perform the rearrangement after the scan instead of performing the rearrangement before the scan.
[0086] Alternatively, instead of a fixed scan order such as a zigzag scan, a horizontal scan, or a vertical scan, the prediction residual rearrangement unit 103b may determine a variable scan order to scan the quantization prediction residuals in ascending order from pixel positions with low probabilities, and output a quantization prediction residual sequence in which the quantization prediction residuals are arranged in ascending order of low probability to the switching unit 103f and the determination unit 103e by performing the scan in the determined scan order.
[0087] The serialization unit 103c performs the same serialization process as the process for a conventional transform skip block on the quantization prediction residuals input from the quantization unit 102b. That is, the serialization unit 103c converts the two-dimensional quantization prediction residuals into a one-dimensional quantization prediction residual sequence in a predefined order (such as a diagonal scan or a zigzag scan), and outputs the quantization prediction residual sequence to the switching unit 103f.
[0088] The accuracy rearrangement unit 103d rearranges the map information input from the evaluation unit 111 into one dimension in the same manner as the prediction residual rearrangement unit 103b based on the accuracy index information input from the sorting unit 103a. Specifically, the accuracy rearrangement unit 103d outputs the map information as a one-dimensional signal in order from the pixel positions with low accuracy (i.e., the pixel positions with low similarity between the reference images). This output signal is herein referred to as the "accuracy column". The accuracy rearrangement unit 103d outputs such an accuracy column to the determination unit 103e.
[0089] The determination unit 103e obtains a similarity representing the correlation between the accuracy column and the quantized prediction residual column from the accuracy column input from the accuracy rearrangement unit 103d and the absolute value of the quantized prediction residual column input from the prediction residual prediction residual rearrangement unit 103b. When the similarity is equal to or greater than the threshold Th (or when it is larger), the determination unit 103e outputs a signal indicating the application of the rearrangement based on the map information to the switching unit 103f. When the similarity is less than (or equal to) the threshold Th, the determination unit 103e outputs a signal indicating the inapplicability of the rearrangement based on the map information to the switching unit 103f.
[0090] Here, Examples 1 to 3 of the similarity obtained from the accuracy column and the quantized prediction residual column are shown below. In the following, the number of pixels in the conversion skip block is N, the accuracy column is pi, the sequence obtained by arranging the accuracy column in reverse order (reverse accuracy column) is qi (i.e., qi = pN-1-i), the quantized prediction residual column is di, and the similarity to be obtained is R. i is an index and takes a range of.
[0091] (1) Example 1 of the similarity used in the determination unit 103e An example using the cosine similarity between the reverse accuracy column and the absolute value of the quantized prediction residual column as the similarity is shown.
[0092]
Equation
[0093] (2) Example 2 of the similarity used in the determination unit 103e As another example, an example of the similarity using the L2 distance is shown.
[0094] [Numerical]
[0095] Here, Qi is a probability sequence obtained by normalizing qi to the range of 0 to 1 (0 indicates the lowest probability and 1 indicates the highest probability). The normalization of the probability is the same as that in the evaluation unit 111. When the probability sequence is originally normalized, Qi = qi. Also, Di is a sequence of numbers obtained by normalizing the absolute value of the quantization prediction residual sequence to the range of 0 to 1. When the maximum possible pixel value of the target image format is minL and the minimum value is maxL, Di can be obtained, for example,
[0096] [Numerical]
[0097] and can be obtained as follows.
[0098] (3) Example 3 of the similarity used in the determination unit 103e As yet another example, an example of the similarity using the L1 distance is shown.
[0099] [Numerical]
[0100] In this way, as the similarity, various similarities and distances that can be calculated between two sequences of numbers can be used. In the case of the similarity (an index where the larger the value, the more similar the two sequences of numbers), the normalized value can be used as the similarity. In the case of the distance (an index where the smaller the value, the more similar the two sequences of numbers), the value obtained by subtracting the normalized distance from 1 can be used as the similarity.
[0101] Note that the threshold Th for comparison with the similarity may be a predetermined constant or a variable value. When it is a variable value, the threshold may be included in the encoded data and transmitted.
[0102] Based on the signal input from the determination unit 103e, the switching unit 103f sends the output of the prediction residual rearrangement unit 103b to the encoding unit 103g when rearrangement based on map information is applicable, and sends the output of the serialization unit 103c to the encoding unit 103g when rearrangement based on map information is not applicable. Note that when rearrangement based on map information is not applicable, the processing of the encoding target block is the same as the conventional conversion skip.
[0103] The encoding unit 103g encodes the quantized prediction residuals in the prediction residual sequence input from the switching unit 103f and outputs encoded data. The encoding unit 103g determines the last significant coefficient included in the quantized prediction residual sequence input from the prediction residual rearrangement unit 103b, and performs encoding on the quantized prediction residual sequence from the beginning to the last significant coefficient. The encoding unit 103g determines in order from the beginning of the quantized prediction residual sequence input from the prediction residual rearrangement unit 103b whether each coefficient is a significant coefficient, assigns an end flag to the last significant coefficient, and does not encode the quantized prediction residuals after the end flag (i.e., zero coefficients), thereby efficiently encoding the significant coefficients.
[0104] For example, as shown in FIG. 8(C), the encoding unit 103g encodes the last significant coefficient in the quantized prediction residual sequence input from the switching unit 103f, that is, the coordinate position (X = 1, Y = 2) in FIG. 8(B), as last_sig_coeff_x and y (end flag). Thereafter, the encoding unit 103g encodes whether a significant coefficient exists as sig_coeff_flag in the reverse order of the scan order, that is, in the order from (3, 3) to (0, 0), starting from the position (1, 2) of the last significant coefficient. In sig_coeff_flag, the coordinate position where a significant coefficient exists is indicated by "1", and the coordinate position where no significant coefficient exists is indicated by "0".
[0105] Furthermore, the encoding unit 103g encodes whether the significant coefficient is greater than 1 as coeff_abs_level_greater1_flag, and encodes whether the significant coefficient is greater than 2 as coeff_abs_level_greater2_flag. For those of the significant coefficients that are greater than 2, the encoding unit 103g encodes the value obtained by subtracting 3 from the absolute value of the significant coefficient as coeff_abs_level_remaining, and encodes the flag indicating the positive or negative of the significant coefficient as coeff_sign_flag.
[0106] By such entropy encoding, the more the position of the last significant coefficient becomes the lower right region (the back in the scan order), the larger the values of last_sig_coeff_x and y become, and the amount of Sig_coeff_flag increases. And the amount of generated information will increase by such entropy encoding.
[0107] When applying the rearrangement based on the map information, rearrangement is performed so that the quantization prediction residuals are encoded in order from the pixel positions with low accuracy (that is, the pixel positions with low similarity between the reference images). Thus, the values of Last_sig_coeff_x and y become smaller, and the amount of Sig_coeff_flag decreases. Thereby, the amount of generated information due to entropy encoding can be reduced.
[0108] Furthermore, as in the example shown in FIG. 4, when the assumption that the pixel positions with low similarity between a plurality of reference images have large prediction residuals does not hold, and the above similarity is less than the threshold Th, the rearrangement based on the map information is not applied. That is, the quantization prediction residuals are encoded in the order according to the conventional diagonal scan. Thereby, the deterioration of the encoding efficiency due to applying the rearrangement based on the map information can be avoided.
[0109] <6. Entropy Decoding Unit> FIG. 9 is a diagram showing an example of the configuration of the entropy decoding unit 200A. As shown in FIG. 9, the entropy decoding unit 200A includes a decoding unit 200a, a sorting unit 200b, a prediction residual rearrangement unit 200c, a deserialization unit 200d, a probability sorting unit 200e, a determination unit 200f, and a switching unit 200g.
[0110] The decoding unit 200a decodes the encoded data generated by the image encoding device 1A, obtains a quantized prediction residual sequence (quantized prediction residual) and information on prediction (intra prediction and motion compensation prediction), outputs the quantized prediction residual sequence to the prediction residual rearrangement unit 200c and the deserialization unit 200d, and outputs the information on prediction to the intra prediction unit 203 and the motion compensation prediction unit 206. When the conversion skip flag obtained from the encoded data indicates the application of conversion skip and the information on prediction represents double prediction, the decoding unit 200a may determine to perform rearrangement based on the map information. Thereby, it is not necessary to include in the encoded data a flag indicating that rearrangement based on the map information is to be performed, and an increase in the flag can be suppressed.
[0111] The sorting unit 200b sorts the probabilities Rij in the map information input from the evaluation unit 208 in ascending order. Since the probabilities Rij are two-dimensionally arranged in the map information, the sorting unit 200b serializes the map information by, for example, zigzag scan to obtain a probability sequence. Then, the sorting unit 200b sorts the probabilities Rij in ascending order, and outputs probability index information in which the probability Rij is used as the index i and the index i is associated with the pixel position (the position of the X coordinate and the position of the Y coordinate) to the prediction residual rearrangement unit 200c and the probability sorting unit 200e.
[0112] The prediction residual rearrangement unit 200c performs the reverse process of the rearrangement process performed by the prediction residual rearrangement unit 103b of the image encoding device 1A. For the transform skip decoding target block, the prediction residual rearrangement unit 200c rearranges the quantized prediction residual sequence input from the decoding unit 200a based on the probability index information input from the sorting unit 200b to deserialize it, and outputs the quantized prediction residual arranged in two dimensions to the switching unit 200g.
[0113] The deserialization unit 200d performs the same deserialization process as the process for the conventional transform skip block. That is, the deserialization unit 200d rearranges the quantized prediction residual sequence into a two-dimensional quantized prediction residual in a predetermined order (for example, zigzag scan from top left to bottom right).
[0114] The probability rearrangement unit 200e performs the same process as the probability rearrangement unit 103d of the image encoding device 1 based on the map information input from the evaluation unit 208, and outputs the probability sequence to the determination unit 200f.
[0115] The determination unit 200f performs the same process as the determination unit 103e of the image encoding device 1.
[0116] When the determination unit 200f outputs a signal indicating the application of the rearrangement based on the map information, the switching unit 200g selects the output of the rearrangement unit 200c. When the determination unit 200f outputs a signal indicating the non-application of the rearrangement based on the map information, the switching unit 200g selects the output of the deserialization unit 200d. The switching unit 200g outputs the selected output (two-dimensional quantized prediction residual) to the inverse quantization unit 201a.
[0117] In an embodiment, the prediction residual rearrangement unit 200c, the deserialization unit 200d, and the switching unit 200g constitute a rearrangement unit. When the similarity is determined by the determination unit 200f to be equal to or greater than a threshold value, such a rearrangement unit 200c rearranges and outputs quantization prediction residuals in an order corresponding to the similarity evaluated by the evaluation unit 208. When the similarity is determined by the determination unit 200f to be less than the threshold value, the quantization prediction residuals are rearranged and output in a predefined order.
[0118] <7. Image Encoding Operations> FIG. 10 is a diagram showing a processing flow in the image encoding apparatus 1A according to the embodiment. The image encoding apparatus 1A executes this processing flow when applying a conversion skip mode and motion compensation prediction to an encoding target block.
[0119] As shown in FIG. 10, in step S1101, the motion compensation prediction unit 109 predicts an encoding target block by performing motion compensation prediction using a plurality of reference images, and generates a prediction image corresponding to the encoding target block.
[0120] In step S1102, the evaluation unit 111 evaluates the similarity between a plurality of reference images for each pixel position, and generates map information representing the accuracy (prediction accuracy) of prediction for each pixel position within the encoding target block.
[0121] In step S1103, the subtraction unit 101 calculates a prediction residual representing the difference in pixel units between the encoding target block and the prediction image.
[0122] In step S1104, the quantization unit 102b performs quantization on the prediction residual calculated by the subtraction unit 101 to generate a quantization prediction residual.
[0123] In step S1105, the entropy encoding unit 103A calculates a similarity representing the correlation between the quantization prediction residual and the map information, and determines whether the similarity is equal to or greater than a threshold value.
[0124] In step S1106, when the similarity is greater than or equal to the threshold, the entropy encoding unit 103A encodes the quantization prediction residuals in ascending order of probability based on the map information and outputs the encoded data. When the similarity is less than the threshold, the entropy encoding unit 103A performs encoding in a predefined order and outputs the encoded data.
[0125] In step S1107, the inverse quantization unit 104a restores the prediction residuals by performing inverse quantization on the quantization prediction residuals input from the quantization unit 102b, and generates restored prediction residuals.
[0126] In step S1108, the synthesis unit 105 reconstructs the target block by synthesizing the restored prediction residuals with the prediction image on a pixel-by-pixel basis, and generates a reconstructed block.
[0127] In step S1109, the loop filter 107 performs filter processing on the reconstructed block.
[0128] In step S1110, the frame memory 108 stores the reconstructed block after the filter processing in frame units.
[0129] <8. Image Decoding Operation> FIG. 11 is a diagram showing a processing flow in the image decoding apparatus 2A according to the embodiment. The image decoding apparatus 2A executes this processing flow when applying the conversion skip mode and motion compensation prediction to the target block.
[0130] As shown in FIG. 11, in step S1201, the decoder 200a of the entropy decoding unit 200A decodes the encoded data to obtain motion vector information, and outputs the obtained motion vector information to the motion compensation prediction unit 206.
[0131] In step S1202, the motion compensation prediction unit 206 predicts the decoding target block by performing motion compensation prediction using a plurality of reference images based on the motion vector information, and generates a prediction image corresponding to the decoding target block.
[0132] In step S1203, the evaluation unit 208 calculates the similarity between a plurality of reference images for each pixel position, and generates map information representing the prediction accuracy (prediction precision) of each pixel position within the block.
[0133] In step S1204, the entropy decoding unit 200A decodes the encoded data to obtain a quantized prediction residual sequence. The entropy decoding unit 200A calculates a similarity representing the correlation between the quantized prediction residual sequence and the map information, and determines whether the similarity is greater than or equal to a threshold value. When the similarity is greater than or equal to the threshold value, the entropy decoding unit 200A rearranges the quantized prediction residuals based on the map information, and outputs the quantized prediction residuals arranged in two dimensions to the inverse quantization unit 201a. When the similarity is less than the threshold value, the entropy decoding unit 200A rearranges the quantized prediction residuals in a predefined order, and outputs the quantized prediction residuals arranged in two dimensions to the inverse quantization unit 201a.
[0134] In step S1205, the inverse quantization unit 201a restores the prediction residual by performing inverse quantization on the quantized prediction residual, and generates a restored prediction residual.
[0135] In step S1206, the synthesis unit 202 reconstructs the target block by synthesizing the restored prediction residual with the prediction image on a pixel-by-pixel basis, and generates a reconstructed block.
[0136] In step S1207, the loop filter 204 performs a filtering process on the reconstructed block.
[0137] In step S1208, the frame memory 205 stores and outputs the reconstructed block after the filtering process on a frame-by-frame basis.
[0138] <9. Summary of the Embodiment> In the image encoding device 1A, the evaluation unit 111 evaluates the similarity between a plurality of reference images in terms of pixels, and outputs information on the evaluation result (map information) to the entropy encoding unit 103A. The entropy encoding unit 103A calculates a similarity representing the correlation between the quantized prediction residual and the map information. When the similarity is equal to or greater than a threshold value, based on the map information, the entropy encoding unit 103A encodes the quantized prediction residual input from the quantization unit 102b in order from the pixel positions with low similarity between the reference images. By encoding the quantized prediction residual in order from the pixel positions with low similarity between the reference images, significant coefficients can be preferentially encoded, and an end flag can be given earlier. Thereby, it is possible to perform efficient entropy encoding on the transform skip block, and the encoding efficiency can be improved.
[0139] In the image decoding device 2A, the evaluation unit 208 evaluates the similarity between a plurality of reference images in terms of pixels, and outputs information on the evaluation result (map information) to the entropy decoding unit 200A. The entropy decoding unit 200A decodes the encoded data to obtain a quantized prediction residual in terms of pixels, calculates a similarity representing the correlation between the quantized prediction residual and the map information, and when the similarity is equal to or greater than a threshold value, rearranges and outputs the quantized prediction residual based on the map information. In this way, by rearranging the quantized prediction residual based on the result of the evaluation by the evaluation unit 208, the entropy decoding unit 200A can autonomously rearrange the quantized prediction residual without information specifying the details of the rearrangement being transmitted from the image encoding device. Thereby, since it is not necessary to transmit information specifying the details of the rearrangement from the image encoding device 1, a decrease in encoding efficiency can be avoided.
[0140] Also, when the similarity representing the correlation between the prediction residual and the map information is less than the threshold value, the encoding efficiency can be further improved by performing rearrangement in a predefined order. Specifically, when the assumption that pixel positions with low similarity among a plurality of reference images have large prediction residuals does not hold, and when the similarity is less than the threshold value, the rearrangement based on the map information is not applied, and the quantized prediction residuals are encoded in the order according to the conventional diagonal scan. Thereby, the deterioration of the encoding efficiency due to applying the rearrangement based on the map information can be avoided.
[0141] <10. Modification Example 1> FIG. 12 is a diagram showing an example of overlapping motion compensation according to Modification Example 1 of the embodiment.
[0142] The evaluation unit 111 of the image encoding device 1A and the evaluation unit 208 of the image decoding device 2A may generate an error map as map information by the method shown below and input it to the rearrangement unit 112. When the error map is input to the rearrangement unit 112, the rearrangement unit 112 performs rearrangement processing of the quantized prediction residuals, with the region where the value of the error map is large being the region with low similarity and the region where the value of the error map is small being the region with high similarity.
[0143] If the luminance signals of two reference images (reference destination blocks) used for generating the prediction image in the dual prediction mode are L0[i, j] and L1[i, j] (where [i, j] are the coordinates within the target block), the error map map[i, j] and its maximum value max_map are calculated by the following formula (2).
[0144] map [i,j] = abs (L0 [i,j] - L1 [i,j]) max_map = max (map [i,j]) ···(2)
[0145] When max_map in formula (2) exceeds 6-bit accuracy (exceeds 64), the error map and the maximum value are updated by shift set so that max_map falls within 6-bit accuracy according to the following formula (3).
[0146] max_map = max_map >> shift map [i,j] = map [i,j] >> shift ··· (3)
[0147] <11. Modified Example 2> The motion compensation prediction unit 109 of the image encoding device 1A and the motion compensation prediction unit 206 of the image decoding device 2A may divide a target block (CU) into a plurality of small blocks, apply different motion vectors for each small block, and be able to switch between unidirectional prediction and bidirectional prediction for each small block. In such a case, for a CU that generates a predicted image using both unidirectional prediction and bidirectional prediction, the evaluation unit 111 of the image encoding device 1A and the evaluation unit 208 of the image decoding device 2A may not calculate map information. On the other hand, when generating a predicted image by bidirectional prediction for all small blocks, the evaluation unit 111 of the image encoding device 1A and the evaluation unit 208 of the image decoding device 2A generate map information.
[0148] Also, the motion compensation prediction unit 109 of the image encoding device 1A and the motion compensation prediction unit 206 of the image decoding device 2A may perform overlapped block motion compensation (OBMC) to reduce the discontinuity of the predicted image at block boundaries with different motion vectors. The evaluation unit 111 of the image encoding device 1A and the evaluation unit 208 of the image decoding device 2A may consider the correction of reference pixels by OBMC when generating map information.
[0149] For example, when the prediction mode of the surrounding blocks used for OBMC correction is dual prediction, the evaluation unit 111 of the image encoding apparatus 1A and the evaluation unit 208 of the image decoding apparatus 2A correct the map information using the motion vectors of the reference images (L0 and L1) used for generating the prediction images of the dual prediction of the peripheral blocks for the region of the prediction image affected by the OBMC correction. Specifically, for the block boundary region, when the motion vectors of adjacent blocks are dual prediction, a weighted average is performed according to the positions with the map information of the adjacent blocks. When the adjacent blocks are in the intra mode or in the case of one-way prediction, the map information is not corrected. In the case of FIG. 12, for the upper block boundary, the map information is generated using L0a and L1a, and for the lower region (the region overlapping with the CU), a weighted average is performed with the map information of the CU. Since the prediction modes of the lower, right, and left CUs are one-way prediction, the map information is not corrected for the regions overlapping with those CUs.
[0150] <12. Modified Example 3> In the above-described embodiment, an example of rearranging the prediction residuals in pixel units has been described, but the prediction residuals may be rearranged in pixel group (small block) units. Such small blocks are blocks composed of 4×4 pixels and may be referred to as CG.
[0151] FIG. 13 is a diagram showing the configuration of the image encoding apparatus 1B according to Modified Example 3 of the embodiment. As shown in FIG. 13, the image encoding apparatus 1B includes a rearrangement unit 112 that rearranges the prediction residuals in pixel group (small block) units. The rearrangement unit 112 generates a prediction image using dual prediction for the target block (CU), and when the CU applies the transform skip mode, based on the above-described error map, rearrangement is performed on the prediction residuals of the CU in small block units (4×4).
[0152] FIG. 14 is a diagram showing an example of the operation of the rearrangement unit 112 according to Modified Example 3 of the embodiment.
[0153] As shown in FIG. 14(A), the subtraction unit 101 of the image encoding device 1B calculates a prediction residual corresponding to the CU. In the example of FIG. 14(A), a prediction residual exists in the upper right region within the CU. The upper right region within the CU can be regarded as a region where the similarity between reference images is low and the prediction accuracy (certainty) is low.
[0154] As shown in FIG. 14(B), the evaluation unit 111 or the rearrangement unit 112 of the image encoding device 1B divides the error map into 4×4 CG units and calculates the average value CGmap of the error in the CG unit by the following formula (4).
[0155]
Equation
[0156] Then, the rearrangement unit 112 rearranges the CGs in descending order of the average value CGmap of the error and assigns an index. In other words, the rearrangement unit 112 rearranges the CGs in descending order of the similarity between reference images and assigns an index. In the example of FIG. 14(B), the numbers within each CG represent the indices after rearrangement. Since the average value CGmap of the CG in the upper right region is large, the priority of scanning (encoding) is set high. Next, as shown in FIG. 14(C), the rearrangement unit 112 rearranges the CGs so that they are scanned (encoded) in ascending order of the index. As a result, as shown in FIG. 14(D), the rearrangement unit 112 outputs the prediction residual rearranged in CG units to the quantization unit 102b.
[0157] The quantization unit 102b quantizes the prediction residual input from the rearrangement unit 112 and outputs the quantized quantization prediction residual to the entropy encoding unit 103B. The entropy encoding unit 103B encodes the CGs in descending order of the average value CGmap of the error to generate encoded data.
[0158] The rearrangement unit 112 rearranges the restored prediction residual output from the inverse quantization unit 104a in units of CG so as to return it to the original arrangement order of the CG, and outputs the restored prediction residual rearranged in units of CG to the synthesis unit 105.
[0159] FIG. 15 is a diagram showing the configuration of an image decoding apparatus 2B according to Modification Example 3 of the embodiment. As shown in FIG. 15, the image decoding apparatus 2B includes a rearrangement unit 209 that rearranges the restored prediction residual output from the inverse quantization unit 201a in units of CG.
[0160] The rearrangement unit 209 generates a prediction image using dual prediction for the target block (CU). When the CU applies the transform skip mode, based on the error map described above, rearrangement in units of CG is performed on the prediction residual of the CU.
[0161] Specifically, the evaluation unit 208 or the rearrangement unit 209 of the image decoding apparatus 2B divides the error map into 4×4 CG units, and calculates the average value CGmap of the errors in units of CG by the above formula (5). Then, the rearrangement unit 209 performs the reverse process of the rearrangement process performed by the rearrangement unit 112 of the image decoding apparatus 2A based on the average value CGmap of the errors in units of CG, and outputs the rearranged prediction residual to the synthesis unit 202.
[0162] In this modification example, when performing determination and switching based on the similarity, the average value of the absolute values of the prediction residuals within the pixel group (CG) may be treated as one element of the prediction residual sequence in the above-described embodiment.
[0163] <13. Modification Example 4> In the above-described embodiment, an example of rearranging the prediction residual in units of pixels or small blocks has been described. However, the prediction residual may be rearranged so as to be inverted horizontally or vertically or both horizontally and vertically.
[0164] As shown in FIG. 16, the subtraction unit 101 of the image encoding device 1B calculates a prediction residual corresponding to a target block. In the example of FIG. 16, a prediction residual exists in the upper right region of the target block. Also, the upper right region of the target block can be regarded as a region where the similarity between reference pixels is low and the prediction accuracy (certainty) is low.
[0165] The rearrangement unit 112 in the modification example 4 of the embodiment calculates the centroid of the error map by the following formula (5).
[0166]
Number
[0167] When the centroid (gi, gj) of the calculated error map is located in the upper right region of the map, that is, when the upper left coordinate is (0, 0) and the lower right coordinate is (m, n),
[0168]
Number
[0169] and
[0170]
Number
[0171] if it is the case, the prediction residual is horizontally flipped.
[0172] When the centroid of the error map is located in the lower left region, that is,
[0173]
Number
[0174] and
[0175]
Number
[0176] If it is the case, the prediction residual is vertically inverted.
[0177] When the center of gravity of the error map is located in the lower right region, that is,
[0178]
Number
[0179] and
[0180]
Number
[0181] If it is the case, the prediction residual is horizontally and vertically inverted.
[0182] Note that when the center of gravity of the error map is in the lower right region, the rearrangement unit 112 may be configured to rotate the prediction residual by 180 degrees instead of horizontally and vertically inverting the prediction residual, or may be configured to change the scan order during coefficient coding from top left to bottom right to bottom right to top left.
[0183] Also, for reduction of processing, without calculating the center of gravity of the error map, the position of the maximum value of the error map may be searched, the position of the maximum value may be regarded as the position of the center of gravity, and the above-described inversion processing may be performed.
[0184] The prediction residual subjected to the inversion processing of the prediction residual by the rearrangement unit 112 in Modification Example 4 of the embodiment is output to the quantization unit 102b.
[0185] The quantization unit 102b quantizes the prediction residual input from the rearrangement unit 112 and outputs the quantized quantization prediction residual to the entropy encoding unit 103B. The entropy encoding unit 103B generates encoded data by encoding in the order from the upper left region to the lower right region of the prediction residual.
[0186] Note that the rearrangement unit 112 performs an inversion process on the restored prediction residual output from the inverse quantization unit 104a based on the position of the center of gravity of the error map, and outputs the rearranged restored prediction residual to the synthesis unit 105.
[0187] <14. Other Embodiments> In the above-described embodiment, an example in which the entropy encoding unit 103A reads all of the two-dimensionally arranged prediction residuals in ascending order of uncertainty and performs serialization processing has been described. However, only the upper several of the two-dimensionally arranged prediction residuals in ascending order of uncertainty may be read, and the other prediction residuals may be read in a fixed order determined by the system. Alternatively, for the two-dimensionally arranged prediction residuals, the reading order may be shifted up or down by a predetermined number according to the uncertainty.
[0188] In the above-described embodiment, inter prediction has been mainly described as motion compensation prediction. In inter prediction, a reference image in a frame different from the current frame is used for predicting the target block of the current frame. However, a technique called intra block copy can also be applied as motion compensation prediction. In intra block copy, a reference image in the same frame as the current frame is used for predicting the target block of the current frame.
[0189] A program for causing a computer to execute each process performed by the image encoding devices 1A and 1B and a program for causing a computer to execute each process performed by the image decoding devices 2A and 2B may be provided. Further, the program may be recorded on a computer-readable medium. By using a computer-readable medium, it is possible to install the program on a computer. Here, the computer-readable medium on which the program is recorded may be a non-transitory recording medium. The non-transitory recording medium is not particularly limited, and may be, for example, a recording medium such as a CD-ROM or a DVD-ROM. Further, circuits for executing each process performed by the image encoding devices 1A and 1B may be integrated, and the image encoding devices 1A and 1B may be configured as semiconductor integrated circuits (chip sets, SoCs). Similarly, circuits for executing each process performed by the image decoding devices 2A and 2B may be integrated, and the image decoding devices 2A and 2B may be configured as semiconductor integrated circuits (chip sets, SoCs).
[0190] As described above, the embodiments have been described in detail with reference to the drawings. However, the specific configuration is not limited to the above, and various design changes and the like can be made without departing from the gist.
Description of Reference Numerals
[0191] 1, 1A, 1B: Image encoding device 2, 2A, 2B: Image decoding device 100: Block division unit 101: Subtraction unit 102: Transformation / quantization unit 102a: Transformation unit 102b: Quantization unit 103A, 103B: Entropy encoding unit 103a: Sorting unit 103b: Prediction residue rearrangement unit 103c: Serialization unit 103d: Accuracy rearrangement unit 103e: Determination unit 103f: Switching unit 103g: Encoding unit 104: Inverse quantization and inverse transformation unit 104a: Inverse quantization unit 104b: Inverse transformation unit 105: Synthesis unit 106: Intra prediction unit 107: Loop filter 108: Frame memory 109: Motion compensation prediction unit 110: Switching unit 111: Evaluation unit 111a: Difference calculation unit 111b: Normalization unit 112: Rearrangement unit 200A: Entropy code decoding unit 200a: Decoding unit 200b: Sorting unit 200c: Prediction residual rearrangement unit 200d: Deserialization unit 200e: Probability rearrangement unit 200f: Judgment unit 200g: Switching unit 201: Inverse quantization and inverse transformation unit 201a: Inverse quantization unit 201b: Inverse transformation unit 202: Synthesis unit 203: Intra prediction unit 204: Loop filter 205: Frame memory 206: Motion compensation prediction unit 207: Switching unit 208: Evaluation unit 208a: Difference calculation unit 208b: Normalization unit 209: Rearrangement unit
Claims
1. A motion compensation prediction unit that generates a block of a prediction image corresponding to a target block by performing dual prediction using a plurality of reference images; Only when performing the dual prediction, an evaluation unit that calculates an absolute difference value in pixel units between the plurality of reference images and calculates an evaluation value indicating a similarity degree calculated according to the absolute difference value in pixel units for each small block composed of a plurality of pixels, which is a unit smaller than the block; An acquisition unit that acquires a prediction residual corresponding to the target block; A synthesis unit that synthesizes the acquired prediction residual with the block of the prediction image to reconstruct the target block, and comprising: The evaluation unit compares the evaluation value with a threshold value; A correction device, wherein the comparison result is used to correct a synthesis target of the synthesis unit in units of the small blocks.
2. A program characterized by causing a computer to function as the correction device according to Claim 1.