Encoding device, decoding device, encoding method, decoding method, encoding program, and decoding program

The use of pre- and post-filters to enhance high-frequency components in video coding addresses the issue of block noise, enhancing coding efficiency and image quality in video compression.

JP7742041B2Active Publication Date: 2025-09-19AKUSERU KK
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2022209738
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-05-11
Filing Date
2022-12-27
Publication Date
2025-09-19
Estimated Expiration
2042-12-27

AI Technical Summary

Technical Problem

Existing video coding methods suffer from degradation in coding efficiency due to block noise caused by discontinuities in region references at high compression rates, particularly in motion compensation codecs.

Method used

A coding device employing pre- and post-filters to enhance high-frequency components of both target and reference images, combined with inter-prediction techniques to improve compression efficiency by reducing prediction residuals.

Benefits of technology

Reduces degradation in coding efficiency by effectively handling high-frequency components, leading to improved image quality and compression rates in video coding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007742041000020
    Figure 0007742041000020
  • Figure 0007742041000021
    Figure 0007742041000021
  • Figure 0007742041000022
    Figure 0007742041000022
Patent Text Reader

Abstract

To reduce the degradation of the encoding efficiency generated in encoding of a moving image.SOLUTION: An encoder for encoding a moving image comprises: a first prefilter part which performs prefiltering on an object image; a second prefilter part which performs prefiltering on a reference image; a prediction part which generates a prediction value for the output of the first prefilter part by using the output of the second prefilter part; and an encoding part which encodes a difference between the output of the first prefilter part and a prediction value generated by the prediction part.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an encoding device, a decoding device, an image processing method, and an image processing program. [Background technology]

[0002] In still image distortion compression such as JPEG (Non-Patent Document 1), pixel values ​​are converted in units of blocks, and then each component is quantized to reduce the number of allocated bits, resulting in the problem of quantization (block) noise. As one solution to this problem, lapped (bi)orthogonal transform, the application of which is JPEG2000 (Non-Patent Document 2) and JPEG-XR (Non-Patent Document 3), has been intensively studied. In particular, JPEG-XR is a two-layer transform coding method based on 4x4 pixel unit blocks, and utilizes a mechanism in which a pre-filter in the encoder emphasizes high-frequency components in areas including block boundaries on the image plane, and a post-filter in the decoder reduces quantization noise along with the high-frequency components (Non-Patent Document 4). A method has also been reported in which an existing still image codec is sandwiched between pre- and post-filters to improve coding efficiency (Non-Patent Document 5). [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] William B. Pennebaker and Joan L. Mitchell. JPEG Still Image Data Compression Standard. Kluwer Academic Publishers, USA, 1st edition, 1992. [Non-patent document 2] C. Christopoulos, A. Skodras, and T. Ebrahimi. The jpeg2000 still image coding system: an overview. IEEE Transactions on Consumer Electronics, Vol. 46, No. 4, pp. 1103-1127, 2000. [Non-patent document 3] Frederic Dufaux, Gary J. Sullivan, and Touradj Ebrahimi. The jpeg xr image coding standard [standards in a nutshell]. IEEE Signal Processing Magazine, Vol.26, No.6, pp.195-204,2009. [Non-patent document 4] TD Tran, Jie Liang, and Chengjie Tu. Lapped transform via time-domain pre- and postfiltering. IEEE Transactions on Signal Processing, Vol. 51, No. 6, pp. 1557-1571, 2003. [Non-patent document 5] Kazunori Kobayashi and Hirohisa Yamaguchi. Image Compression Using Prefilters and Postfilters. Proceedings of the IEICE General Conference, 1997. Information Systems, No. 2, p. 28, Mar 1997. Summary of the Invention [Problem to be solved by the invention]

[0004] As mentioned above, there is a known method for improving the coding efficiency of still images by sandwiching the codec between a pre-filter and a post-filter. In video coding using inter-coding and other methods, the introduction of pre- and post-filtering can also be expected to improve coding efficiency. Inter-prediction is a technique that increases the compression rate of video data by generating a predicted image from frames (reference images) before and after the frame to be coded (target image), detecting the difference with the target image (prediction residual), and encoding it. Inter-prediction uses motion compensation to further improve compression efficiency. Motion compensation encodes the range, direction, and speed (motion vector) of a moving object (the same object that appears in the previous and next frames) within the video, and reflects this in the creation of a predicted image. In general, motion compensation codecs have a problem in that discontinuities in region references appear as block noise at high compression rates. An object of one aspect of the present invention is to reduce degradation in coding efficiency that occurs in coding of moving images. [Means for solving the problem]

[0005] The present invention provides a coding device for coding moving images, a first prefilter unit that performs prefiltering on the target image and generates blocks of a prefiltered target image; a second prefilter unit that performs prefiltering on a reference image and generates blocks of a prefiltered reference image; a first prediction unit that predicts pixel values ​​of blocks of the target image that are not prefiltered; a second prediction unit that predicts pixel values ​​of the prefiltered target image; and a coding unit that encodes the target image and generates coded data, wherein the second prediction unit performs intra prediction, inter prediction, or hierarchical inter prediction using an output of the first prefilter unit or outputs of the first prefilter unit and the second prefilter unit. It is characterized by: [Effects of the Invention]

[0006] According to the present invention, it is possible to reduce the degradation of coding efficiency that occurs in coding of moving images. [Brief explanation of the drawings]

[0007] [Figure 1] 1 is a diagram showing a schematic configuration of an image processing system according to an embodiment of the present invention; [Figure 2] FIG. 1 is a diagram illustrating the functional configuration of an image processing device serving as an encoder according to a first embodiment of the present invention. [Figure 3] FIG. 1 is a diagram illustrating the functional configuration of an image processing device serving as a decoder according to a first embodiment of the present invention. [Figure 4] FIG. 2 is a diagram illustrating a processing flow in the encoder 10. [Figure 5] FIG. 2 is a diagram illustrating a processing flow in the decoder 30. [Figure 6]FIG. 10 is a diagram illustrating the functional configuration of an image processing device serving as an encoder according to a second embodiment of the present invention. [Figure 7] FIG. 10 is a diagram illustrating the functional configuration of an image processing device serving as a decoder according to a second embodiment of the present invention. [Figure 8] 10 is a diagram for explaining operations on image data (blocks) in the image processing system of the second embodiment. FIG. [Figure 9] FIG. 10 is a diagram showing a filter block starting from an even pixel. [Figure 10] FIG. 10 is a diagram showing a filter block starting from an odd pixel. [Figure 11] FIG. 10 is a diagram illustrating pre-filtering. [Figure 12] FIG. 10 is a diagram illustrating pre-filtering for a filter block starting from an odd pixel. [Figure 13] FIG. 10 is a diagram illustrating a boundary pattern of post-filtering. [Figure 14] FIG. 10 is a diagram showing image data used in an experiment. [Figure 15] FIG. 10 is a graph showing the measurement results. [Figure 16] FIG. 10 is a graph showing the measurement results. [Figure 17] FIG. 10 is a graph showing the measurement results. [Figure 18] FIG. 10 is a graph showing the measurement results. [Figure 19] FIG. 10 is a graph showing the measurement results. [Figure 20] 10 is a flowchart illustrating an encoding process performed by an encoder. [Figure 21] 10 is a flowchart illustrating a decoding process executed by a decoder. [Figure 22] FIG. 1 is a block diagram illustrating an embodiment of a computer device. DETAILED DESCRIPTION OF THE INVENTION

[0008] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. FIG. 1 is a diagram showing a schematic configuration of an image processing system according to this embodiment. The image processing system 1 shown in FIG. 1 includes an image processing device 10 that functions as an encoder for compressing (encoding) an image, and an image processing device 30 that functions as a decoder for restoring (decoding) an image. An encoder is also called a coding device, and a decoder is also called a decoding device. The encoder 10 compresses and encodes the target image to generate encoded data. The decoder 30 decodes the coded data created by the encoder 10 to restore and display the target image. The encoder 10 and decoder 30 can be configured as a computer device 100 as described below, but may also be configured as an image processing device (module) incorporated into a computer device or other appliance product. The encoder 10 is incorporated into a computer device or the like that creates images, and converts the created images into coded data that can be decoded by the decoder 30 . The decoder 30 is configured as a VDP (Video Display Processor) incorporated into equipment that displays images, such as a gaming machine or signage, and decodes the coded data created by the encoder 10 and displays it on a liquid crystal display device or the like. In this specification, an image includes a moving image made up of a plurality of frames. The image processing system 1 of this embodiment aims to prevent a deterioration in coding efficiency when video is coded using inter (inter-frame) prediction. In the following description, a frame of a video to be coded is referred to as a target image, and a frame referenced in inter-frame prediction of this target image is referred to as a reference image. Note that the reference image can be a frame immediately before or after the target image, or a frame that is not immediately before or after the target image.

[0009] <Embodiment 1> FIG. 2 is a diagram illustrating the functional configuration of the image processing device as an encoder according to the first embodiment of the present invention. The encoder 10 includes a control unit 11, a storage unit 12, a communication unit 13, and an input unit 14. The control unit 11 includes a receiving unit 20, a predicting unit 60, a first pre-filter unit 24A, a second pre-filter unit 24B, an encoding unit 25, and an output unit . At least one of the processing units included in the control unit 11 may be realized by hardware. The receiving unit 20 receives, via the communication unit 13 and the input unit 14, an image (video image) to be encoded. The first pre-filter unit 24A performs predetermined filtering on the entire target image, and the second pre-filter unit 24B performs predetermined filtering on the entire reference image. Here, the predetermined filtering is intended to increase the flatness of the frequency spectrum of the target image. The frequency components contained in the image signal of a natural image are not constant across frequency bands, and the energy (power and amplitude) of signals in high frequency bands tends to be smaller than that in low frequency bands. Therefore, in such cases, processing is included to emphasize the high frequency components of the pixels of the target image and reference image. The prediction unit 60 performs inter prediction on the image data filtered by the first pre-filter unit 24 A. Inter prediction calculates predicted values ​​(predicted image) of pixel values ​​of a target image by referring to a reference image.

[0010] Inter-prediction is a technique for generating a predicted image of a target frame (target image) from frames (reference images) before and after the target frame. In this embodiment, the difference (prediction residual) between the predicted image and the target image is coded across the entire target frame. This significantly improves coding efficiency (compression efficiency). In inter-prediction, motion compensation is used to further improve compression efficiency. Motion compensation encodes the range, direction of movement, and speed (motion vector) of a moving object (the same object present in previous and next frames) within a video sequence, and reflects this in the creation of a predicted image. Due to camera shake during shooting, even if the objects depicted in the video sequence are nearly identical, there is a possibility that a huge difference will be generated when calculating the prediction residual. To avoid this, the direction and speed of movement (motion vector) across all previous and next frames can be coded and included in the frame data.

[0011] FIG. 3 is a diagram illustrating the functional configuration of the image processing device as a decoder according to the first embodiment of the present invention. The decoder 30 includes a control unit 31, a storage unit 32, a communication unit 33, and an input unit . The control unit 31 includes a receiving unit 40 , a predicting unit 62 , a third pre-filter unit 44 , a post-filter unit 45 , a decoding unit 46 , and an output unit 47 . At least one of the processing units included in the control unit 31 may be realized by hardware. The receiving unit 40 receives encoded data to be decoded via the communication unit 33 and the input unit 34 . The third pre-filter unit 44 performs pre-filtering on the previous or next frame that will be used as a reference image during decoding. The decoding unit 46 decodes the prediction residual (encoded data) received by the receiving unit 40. The prediction unit 62 adds the pre-filtered reference image and the prediction residual output from the decoding unit 46 together to perform inter prediction. The post-filter unit 45 performs post-filtering on the output of the prediction unit 62 . The output unit 47 outputs the image processing result to the outside of the decoder 30. If the decoder 30 is an image processing device such as a VDP, the output by the output unit 47 may be the display of the decoded image on a display device.

[0012] The prefiltering process in the third prefilter unit 44 increases the flatness of the frequency spectrum of the reference image, similar to the first prefilter unit 24A and second prefilter unit 24B of the encoder 10 described above, and if the first and second prefilter units 24A and 24B perform processing that emphasizes high-frequency components, the third prefilter unit 44 also performs filtering processing that emphasizes high-frequency components. The postfilter unit 45 also performs filtering processing on the output of the prediction unit 62, and the filtering process in the postfilter unit 45 is processing that is the opposite of the processing performed in the first to third prefilter units, i.e., if the first to third prefilter units perform processing that emphasizes high-frequency components, the postfilter unit 45 performs processing to weaken the high-frequency components. The storage unit 32 can store temporary files and temporary data used in image processing calculations, and output data after image processing. The storage unit 32 can also store coded data (DC components and prediction residuals) to be decoded, which is created by the encoder 10. The communication unit 33 connects the decoder 30 to a network, enabling communication with external devices.

[0013] FIG. 4 is a diagram for explaining the processing flow in the encoder 10, and FIG. 5 is a diagram for explaining the processing flow in the decoder 30. As shown in FIG. In the encoder 10, pixels for one frame of the video to be encoded are supplied to the first pre-filter unit 24A, which then performs a process to increase the flatness of the frequency spectrum of the input data, for example, a filtering process to emphasize high-frequency components, and the pixel values ​​with the emphasized high-frequency components are supplied to the prediction unit 60. In addition, pixels for one frame of a reference image, which is the frame before or after the target image, are supplied to the second pre-filter unit 24B, and the second pre-filter unit 24B also performs a process to increase the flatness of the frequency spectrum of the data, for example, a filtering process that emphasizes high-frequency components, and the pixel values ​​with emphasized high-frequency components are supplied to the prediction unit 60. The prediction unit 60 calculates a predicted value (predicted image) using inter-prediction based on the input filtered target pixel group and reference pixel group, and supplies the prediction difference (residual) to the encoding unit 25. The encoding unit 25 encodes the residual and supplies it to the output unit 26. The prediction difference (residual) is the difference between the target image and the predicted image, and is calculated by subtraction. The prediction residual can also be quantized, or the quality and code amount can be increased or decreased by adjusting the quantization parameter. Furthermore, the prediction residual can be expressed as frequency components by performing frequency transform, or the prediction difference can be expressed by prediction using neighboring pixels. The encoding unit 25 encodes the prediction residual. Entropy coding can also be used for encoding. In this embodiment, an example has been given in which the prediction unit 60 performs subtraction processing between the target image and the predicted image in addition to prediction processing including the above-mentioned prediction processing using motion vectors and level adjustment of predicted pixels, but it is also possible for the prediction unit 60 to only perform prediction processing, and for the subtraction processing between the target image and the predicted image to be performed in a separately provided processing unit, and the output of that processing unit may be supplied to the next-stage encoding unit 25, or the encoding unit 25 may perform subtraction processing between the target image and the predicted image. In this way, in inter prediction, only the data of the prediction difference between the target image and the predicted image is coded, so the amount of data is small and a high compression rate can be achieved.

[0014] Meanwhile, in the decoder 30, the decoding unit 46 decodes the input code to obtain a prediction difference (residual), and supplies the prediction difference to the prediction unit 62. A group of pixels for one frame of a reference image, which is a frame before or after a target image that is the target of inter prediction, is supplied to the third pre-filter unit 44, which performs a filtering process to enhance the flatness of the frequency spectrum. For example, if the first pre-filter unit and the second pre-filter unit perform a filtering process to emphasize high-frequency components as described above, the third pre-filter unit 44 also performs a filtering process to emphasize high-frequency components, and the prediction unit 62 adds the pixel value(s) in which high-frequency components have been emphasized output from the third pre-filter unit 44 to the prediction difference supplied from the decoding unit 46, and outputs pixel values ​​in which high-frequency components have been emphasized to the post-filter 45. The post-filter 45 performs a filtering process to weaken the high-frequency components on the input group of pixel values ​​in which high-frequency components have been emphasized, and outputs the decoded pixel values ​​to the output unit 47. The prediction residual is decoded by performing the reverse of the encoding, such as inverse quantization, inverse frequency transformation, prediction using neighboring pixels, or decoding symbols for entropy codes. Note that, as with the above-described encoder, in this embodiment, an example has been described in which the prediction unit 62 performs an addition process between a target image and a predicted image in addition to the prediction process using the above-described motion vector and prediction process including level adjustment of predicted pixels, but it is also possible for the prediction unit 62 to only perform the prediction process, and for the addition process between the target image and the predicted image to be performed by a separately provided processing unit, with the output of that processing unit being supplied to the next-stage post-filter 45, or for the post-filter 45 to perform the addition process between the target image and the predicted image and then perform filtering on the data after the addition process. In this way, it is sufficient to determine which component is used to obtain the prediction residual and which component is used to obtain the addition of the predicted pixel value to the prediction residual, as this applies to embodiment 2 described below.

[0015] 15 to 19 show the measurement results of the coding efficiency when pre-filtering is performed and when it is not performed, where FIG. 15 shows the measurement results using data (1) described later, FIG. 16 shows the measurement results using data (2) described later, FIG. 17 shows the measurement results using data (3) described later, FIG. 18 shows the measurement results using data (4) described later, and FIG. 19 shows the measurement results using data (5) described later. The vertical axis in Figures 15 to 19 is PSNR ([dB]), with a higher value indicating less distortion and better performance. The horizontal axis is the compression ratio, with a lower value indicating better performance, with less coding required to describe the data compared to the original data. In general, the closer the data is plotted to the upper left of the graph, the better the codec. It can be seen that in scenes where the subject in data (1) to (3) has moved or deformed and the correlation is low, the coding efficiency is improved even when filtering is applied to the entire image. Details of the filtering process will be described in the second embodiment below.

[0016] In the encoding and decoding of moving images using inter-prediction as in this embodiment, high-quality image signals can be obtained by performing filtering processing, for example, to emphasize high-frequency components on both the target image and the reference image. This is presumably because the frequency components contained in the image signal of a natural image are not constant across frequency bands, and the energy of signals in high-frequency bands tends to be smaller than that in low-frequency bands. When processing such an image signal that causes uniform distortion across the entire frequency band, such as encoding such as distortion data compression, is performed, image signals in high-frequency bands with low signal component energy are more susceptible to noise caused by signal distortion. This reduces the S / N ratio in such frequency bands, which is one of the causes of image quality degradation. Therefore, it is conceivable that, before encoding, a pre-filter is used to flatten the frequency spectrum of the image signal, and the level of the signal in the low-energy frequency band is raised before performing processing such as encoding, transmission, and decoding, and after decoding, a post-filter is used to attenuate the amplified frequency band of the image signal and return it to its original state, thereby reducing the effects of noise. Note that pre-filtering amplifies signals in low-energy frequency bands, which increases the overall signal energy and residual error, tending to increase the amount of code. However, the improvement in image quality outweighs the increase in the amount of code, allowing for efficient coding.

[0017] In the present embodiment, the reference image is described as a frame before or after the target image, but is not limited to a frame immediately before or after the target image. Also, while a pre-filter that emphasizes high-frequency components has been described as an example, high-quality moving images can be obtained by applying a filter process to the target image and reference image that increases the amount of energy in frequency bands with low signal energy according to the frequency band characteristics of the target moving image, and then encoding and decoding the resultant images. Furthermore, the filter characteristics of the first to third pre-filter units 24A, 24B, and 44 may be fixed, or may be variable so as to flatten the frequency spectrum of the input image signal, and information about the filter characteristics may be added to the data encoded by the encoder and transmitted, and the decoder may control the filter characteristics of the third pre-filter unit 44 and the post-filter unit 45 based on the information about the filter characteristics.

[0018] <Embodiment 2> Next, a second embodiment of the present invention will be described. In the first embodiment described above, as is clear from the experimental results shown in Fig. 15 and Fig. 16, when a video having a high correlation between a target image and a reference image is coded, if filtering that emphasizes high frequency components uniformly is performed on the entire image, the coding efficiency tends to decrease compared to when no filtering is performed. Therefore, in the second embodiment, instead of performing pre-filtering on the entire video, the correlation between the target image and the reference image is determined, and pre-filtering is not performed on areas with high correlation, but on areas with low correlation.

[0019] 6 is a diagram illustrating the functional configuration of an image processing device as an encoder according to a second embodiment of the present invention. Note that the same components as those in the first embodiment are given the same reference numerals, and repeated explanations will be omitted. The encoder 10 includes a control unit 11, a storage unit 12, a communication unit 13, and an input unit 14. The control unit 11 includes a receiving unit 20, a dividing unit 21, a first prediction unit 22A, a second prediction unit 22B, a first mode determination unit 23, a first pre-filter unit 24A, a second pre-filter unit 24B, an encoding unit 25, and an output unit 26. At least one of the processing units included in the control unit 11 may be realized by hardware. The receiving unit 20 receives, via the communication unit 13 and the input unit 14, an image (video image) to be encoded. The division unit 21 divides a video frame (target image) into blocks of 8×8 pixels, 16×16 pixels, etc. When performing inter prediction, the division unit 21 also divides a reference image into blocks. The encoder 10 encodes the target image in units of blocks. The first prediction unit 22A performs prediction on the blocks of the target image divided by the division unit 21. The second prediction unit 22B performs prediction on image data (filtered blocks) obtained by filtering the blocks of the target image divided by the division unit 21 using the first pre-filter unit 24A. The prediction modes used in the first prediction unit 22A and the second prediction unit 22B can include intra prediction, inter prediction, hierarchical inter prediction, and the like.

[0020] The first mode determination unit 23 determines the prediction mode to be used for the block to be coded and the coding mode for whether or not to perform pre-filtering. Multiple prediction modes, such as intra-prediction, inter-prediction, and hierarchical inter-prediction, can be performed, and the prediction mode that produces the smallest prediction residual can be selected. Whether or not to perform pre-filtering can be determined based on whether or not the prediction residual is equal to or greater than a predetermined value. The prediction mode determination is not limited to this, and frequency analysis, the magnitude relationship of feature quantities such as autocorrelation functions and cross-correlation functions, and threshold determination can also be used. The first pre-filter unit 24A performs pre-filtering on blocks of the target image. The second pre-filter unit 24B performs pre-filtering on blocks of a reference image used when the second predictor 22B performs inter-prediction. The encoding unit 25 encodes the prediction residual in the prediction mode selected by the first mode determination unit 23 to generate encoded data of the target image. The prediction residual can also be quantized. The quality and code amount can be increased or decreased by adjusting the quantization parameter. Encoding can also be performed by performing frequency transformation on the prediction residual or prediction using neighboring pixels, or by entropy coding the symbol to be encoded. The output unit 26 outputs the image processing result (encoded data) to the outside of the encoder 10. The storage unit 12 can store temporary files and temporary data used in image processing calculations, and output data after image processing. The communication unit 13 connects the encoder 10 to a network, enabling communication with external devices.

[0021] FIG. 7 is a diagram illustrating the functional configuration of an image processing device serving as a decoder according to the second embodiment of the present invention. The decoder 30 includes a control unit 31, a storage unit 32, a communication unit 33, and an input unit . The control unit 31 includes a reception unit 40, a combination unit 41, a third prediction unit 42A, a fourth prediction unit 42B, a second mode determination unit 43, a third pre-filter unit 44, a post-filter unit 45, a decoding unit 46, and an output unit 47. At least one of the processing units included in the control unit 31 may be realized by hardware. The receiving unit 40 receives encoded data to be decoded via the communication unit 33 and the input unit 34 . The second mode determination unit 43 determines the prediction mode used for a block in the coded data and whether pre-filtering has been performed. The third prediction unit 42A performs prediction for a block determined by the second mode determination unit 43 to be pre-filtered. The fourth predictor 42B performs prediction on a block determined by the second mode determiner 43 to have undergone pre-filtering. When the second mode determination unit 43 determines that the encoding mode of the block to be decoded is prefiltered and inter-prediction, the third prefilter unit 44 performs prefiltering on blocks of the previous or next frame that will be used as reference images during decoding.

[0022] The decoding unit 46 decodes the prediction residual and adds it to the predicted value predicted by the third prediction unit 42A or the fourth prediction unit 42B to obtain a decoded image. The prediction residual is decoded by performing the reverse of the encoding, such as inverse quantization, inverse frequency transform, or prediction using adjacent pixels, or decoding symbols for entropy codes. The post-filter unit 45 performs post-filtering on the blocks that are determined by the second mode determination unit 43 to require pre-filtering. The combining unit 41 combines the decoded blocks to restore the target image (original image). The output unit 47 outputs the image processing result to the outside of the decoder 30. If the decoder 30 is an image processing device such as a VDP, the output by the output unit 47 may be the display of the decoded image on a display device. The storage unit 32 can store temporary files and temporary data used in image processing calculations, and output data after image processing. The storage unit 32 can also store coded data to be decoded (DC components and prediction residuals for each block, which will be described later) created by the encoder 10. The communication unit 33 connects the decoder 30 to a network, enabling communication with external devices.

[0023] FIG. 8 is a diagram for explaining the operation of image data (blocks) in the image processing system of the second embodiment. The processes executed by the encoder and decoder included in the image processing system of the second embodiment will be described with reference to FIGS. The image compression (encoding) and restoration (decoding) performed by the image processing device of this embodiment will be described below. As explained in the first embodiment, in a video with little subject movement overall and high correlation, performing pre-filtering to emphasize high frequencies on all regions of the target image and reference image will degrade the coding efficiency. This is because the image regions include regions with high correlation and regions with low correlation, and pre-filtering that emphasizes high frequencies is effective for regions with low correlation but not for regions with high correlation. As shown in the experimental results below, coding efficiency improves when pre-filtering is not applied to regions with high correlation, and pre-filtering is applied only to regions with low correlation.

[0024] The encoder 10 of the second embodiment divides the target image into blocks 50 (for example, 16×16 pixels or 8×8 pixels) in advance. The encoder 10 uses a first prediction unit 22A to perform prediction on each block using several methods, such as intra prediction, inter prediction, and hierarchical inter prediction, and evaluates the prediction residual. The second prediction unit 22B performs prediction on each block using several methods, such as intra prediction, inter prediction, and hierarchical inter prediction, on data prefiltered by the first prefilter unit 24A, and evaluates the residual. When predicting a block of a target image prefiltered by the second prediction unit 22B using inter prediction or hierarchical inter prediction, a reference image prefiltered by the second prefilter unit 24B is used. The first mode determination unit 23 of the encoder 10 determines the prediction mode with the smallest prediction residual and whether prefiltering is performed in that prediction mode, and encodes the pixels of the block 50 in the determined encoding mode. In addition, instead of an algorithm that performs intra prediction, inter prediction, and hierarchical inter prediction on blocks that do not undergo pre-filtering and blocks that have undergone pre-filtering, and then determines the prediction mode with the smallest prediction residual, an algorithm may be adopted that first performs prediction on blocks that do not undergo pre-filtering, and if the prediction residual is greater than or equal to a predetermined value in any prediction mode, performs pre-filtering on blocks of the target image to perform intra prediction, or performs pre-filtering on blocks of the target image and blocks of the reference image to perform inter prediction or hierarchical inter prediction, as in the embodiment described below. For example, if the prediction residual from intra-prediction or inter-prediction is greater than or equal to a predetermined value, it indicates poor prediction accuracy in intra-prediction or inter-prediction (i.e., low correlation with other blocks in the same frame or low correlation with blocks in other frames). For such a block 50, the encoder 10 applies a prefilter to the block in the target image using the first prefilter unit 24A. The encoder 10 then performs predictions on each prefiltered block using several methods, including inter-prediction, to evaluate the prediction residual. For inter-prediction of a prefiltered block, the encoder 10 also applies a prefilter to the block in the reference image using the second prefilter unit 24B, and recalculates motion compensation for the filtered reference image. The encoder 10 encodes the pixel values ​​of the prefiltered block 50 in the target image using a method that minimizes the prediction residual. Therefore, an image with a large prediction residual from inter-prediction and to which a prefilter has been applied can be encoded using inter-prediction or intra-prediction using the pre-filtered reference image.

[0025] The encoder 10 of this embodiment encodes blocks 50 whose prediction residuals from any prediction mode, including inter-prediction, satisfy a condition (are less than a predetermined value) on a block-by-block basis, while applying a pre-filter to blocks 50 that do not satisfy the condition and then encoding them on a block-by-block basis. In this embodiment, (1) Block prediction without pre-filtering (intra prediction, inter prediction, etc.) (2) For blocks with poor prediction accuracy, prediction in small block units with pre-filtering (intra prediction, inter prediction, etc.) Video is coded by performing two-stage prediction. The coded data of a moving image coded by this method contains a mixture of regions (blocks) with low correlation that have been prefiltered and regions with high correlation that have not been prefiltered. Furthermore, pre-filtering is performed during encoding, and corresponding post-filtering is performed during decoding. However, if there is a mixture of areas where pre-filtering has been performed and areas where pre-filtering has not been performed, the post-filter on the boundary between these areas will result in different inverse transformations depending on the mixture pattern.

[0026] This will be explained in more detail with reference to FIG. As shown in FIG. 8(A), the dividing unit 21 of the encoder 10 divides the image to be encoded into blocks 50. The first prediction unit 22A of the encoder 10 performs prediction using several methods, including inter prediction, on the selected block 50. In this embodiment, intra prediction is used as a prediction other than inter prediction. An example of intra prediction is a method of using the average value of decoded blocks located above and to the left of block 50 as a predicted value. It is also possible to use a predicted value corrected by an offset δ so that the predicted value is closer to the average value of the block 50. The offset δ is coded separately. On the other hand, one example of inter prediction is a motion-compensated prediction method in which a reference region block corresponding to the current block 50 is searched for in a reference image, and pixel values ​​of the reference region block corresponding to each pixel constituting the current block 50 are used as predicted values. Furthermore, when performing inter prediction using a prefilter, prefiltering is performed on the current block 50 and the reference region, and motion compensation is performed using the same procedure as in normal inter prediction. Inter / intra prediction can be performed uniformly on the current block 50, or the block 50 can be re-divided and predicted hierarchically. When performing hierarchical prediction, it is also possible to apply different prediction methods to each of the divided blocks. The first mode determination unit 23 of the encoder 10 determines a coding mode that is expected to increase the coding efficiency of the block 50. The coding mode includes the selection of a prediction mode and the presence or absence of a pre-filter. The prediction mode can be selected from the above-mentioned intra prediction, inter prediction, hierarchical prediction methods, etc. The prediction mode can be determined based on the size of the prediction residual and the size of the image distortion. Since the smaller the prediction residual, the smaller the code amount tends to be, the prediction mode can be selected by evaluating the prediction residual without directly evaluating the code amount.

[0027] In the second embodiment, efficient coding and prediction method determination are performed by performing the following hierarchical inter-coding. In the example of FIG. 8, the block 50 is 16×16 pixels, but motion compensation may also be applied to blocks of, for example, 8×8 pixels. When using hierarchical inter-coding, inter-prediction and intra-prediction are first performed on the block 50. If the prediction residual of either the inter-prediction or intra-prediction on the block satisfies a predetermined condition depending on the quantization coefficient Q that determines the image quality, the processing ends without dividing the block 50. At this time, if the prediction residual from inter-prediction is small, the prediction mode is determined to be inter-coding, and the prediction residual is not coded. If the prediction residual from intra-prediction is small, the prediction mode is determined to be intra-prediction, and the offset δ is coded, but the prediction residual is not coded. Motion compensation is recalculated for blocks that do not satisfy the condition. Furthermore, when prediction is performed for a block, if two or more of the four divided regions (an 8x8 pixel block in the example of Figure 8(B)) do not satisfy the condition, or if two regions that satisfy the condition are located diagonally within the block, recalculation is performed for each of the four divided regions. From this point on, the prediction residual is evaluated in 4x4 pixel and 2x2 pixel units, and the prediction mode for each region is determined hierarchically. The prediction mode for 4x4 pixel and 2x2 pixel blocks can be determined by evaluating the prediction residual of inter prediction or intra prediction, just as with 16x16 pixel and 8x8 pixel blocks. If the prediction residual satisfies a predetermined condition, the prediction residual is not coded, but is coded using inter prediction or intra prediction. If the predetermined condition is not met, the region is re-divided and coding is performed recursively. Finally, for the 2x2 pixel units that do not satisfy the condition, either inter prediction or intra prediction is applied, and entropy coding is performed on the Hadamard transform coefficients of the residual. Specifically, the predicted value p(·) of each pixel value {a, b, c, d} using the processed pixel values ​​{x, y, z, w} of the top left two pixels of the 2 × 2 block shown in Figure 8 (C) is p(a)={x+z+2f(a)} / 4 p(b)={y+f(b)} / 2 p(c)={w+f(c)} / 2 p(d)=f(d) Here, f(·) is the corresponding pixel value of the predicted block in inter prediction, and the average pixel value of the corresponding 2×2 block of the current block in intra prediction. In this way, a large block may be divided into smaller blocks of pixels little by little, and inter prediction and intra prediction may be performed to evaluate the prediction residual, but block 50 may also be divided directly into 2×2 small blocks.

[0028] As shown in Figure 8 (B-3), for a block 50 in which the prediction residual of intra prediction is less than a predetermined value, the first mode determination unit 23 of the encoder 10 determines the coding method for that block 50 to be intra coding and performs coding using intra prediction. As shown in Figure 8 (B-2), for a block 50 in which the prediction residual of intra prediction is equal to or greater than a predetermined value but the prediction residual of inter prediction is less than a predetermined value, the first mode determination unit 23 of the encoder 10 determines the coding method for that block 50 to be inter coding and performs coding using inter prediction. As shown in FIG. 8(B-1), for a block 50 in which both the prediction residual of intra prediction and the prediction residual of inter prediction are equal to or greater than a predetermined value, the first pre-filter unit 24A of the encoder 10 performs pre-filtering on such a block 50. The block on which prefiltering has been performed is divided into four regions that are half the size vertically and horizontally, and a prediction mode is determined for each divided region. If the prediction mode cannot be determined, that is, if the prediction residual of the divided region does not satisfy a predetermined condition, division is repeated until the region size becomes 2 x 2, and prediction mode determination is recursively performed. When inter prediction is performed, prefiltering is also performed on the reference image, and the encoding unit 25 of the encoder 10 finally encodes the information of each block. Here, pre-filtering will be described for regions (8×8 pixel blocks) obtained by dividing the 16×16 pixel block 50 described in FIG. 8 into four, with reference to FIGS. 9 and 10. FIG.

[0029] 9 is a diagram showing a filter block starting from an even pixel of a block. TB is a block of 2×2 pixels, which is the minimum size that can be generated by hierarchical inter-coding. When an even pixel is used as the starting point, the block TB itself becomes the filter block FB. FIG. 10 is a diagram showing filter blocks starting from odd-numbered pixels of a small block. The encoder 10 can define a filter block FB that straddles the boundary of a block TB. In this embodiment, a filter block of 2×2 pixels is described, but the number of pixels in the block TB may be increased to define a filter block FB of 4×4 pixels. The encoder 10 adds a pre-filter (described later) to the filter block FB. The pre-filter is a filter for emphasizing high frequency components at the boundary between blocks. Specifically, the encoder 10 applies 2×2 pixel filter blocks FB1, FB2, FB3, etc. to the boundaries of blocks TB1, TB2, TB3, etc. Like block TB, filter block FB also contains four pixels, so each pixel in block TB is filtered by a different filter block FB. In other words, by filtering the pixels in block TB with different filter blocks FB rather than a single filter block FB, noise at the boundaries of block TB can be reduced. Note that filter blocks FB1, FB2, FB3, etc. may extend beyond block 50. In such cases, the pixels of the filter blocks at positions extending beyond block 50 are compensated for with appropriate pixel values, such as the luminance values ​​of pixels adjacent to that position. If filtering is performed starting only from even vertices as in Figure 9, the block boundary and the filtering boundary will coincide when the motion vector in motion compensation prediction is an even unit. In this case, noise that occurs at the block boundary on the decoder side cannot be reduced by the post-filter. In contrast, by additionally performing pre-filtering starting from odd pixels as in Figure 10, smoothing by post-filtering can be applied to the block boundary even when the motion vector is odd, making it possible to reduce noise. Note that either filtering starting from even vertices or filtering starting from odd vertices can be performed first.

[0030] FIG. 11 is a diagram illustrating pre-filtering. The filter block FB shown in FIGS. 9 and 10 is made up of 2×2 pixels. For example, two pixels p -1 , p1 will be used as an example to explain pre-filtering. As shown in FIG. -1 , p1 are filtered by the following operation to obtain pixel p -1 , p1 is weighted, and the filtered pixel p'-1 , p'1 is calculated. TIFF0007742041000001.tif1057Here, λ∈[1,2]. The pre-filter is a filter for emphasizing high frequency components at the boundaries of blocks, and λ is a coefficient that determines the strength of the pre-filter's emphasis on high frequencies.

[0031] FIG. 12 is a diagram showing pre-filtering for a filter block starting from an odd pixel. The encoder 10 applies pre-filtering to four pixels in the filter block FB in block units. At this time, the encoder 10 performs the pre-filtering described with reference to Fig. 11 two-dimensionally (vertically and horizontally). The encoder 10 performs two-dimensional pre-filtering in units of blocks using a transformation matrix (pre-filter) F as a pre-filter. The encoder 10 performs two-dimensional pre-filtering in units of blocks using four pixels (p -1,-1 , p -1,-1 , p 1,-1、 p 1,1 ) and apply a pre-filter F to it. Calculate TIFF0007742041000002.tif2445. The pre-filter F is The file is TIFF0007742041000003.tif19110. Such a pre-filter F is used for four pixels (2×2) as a one-dimensional filter shown in FIG. This is a two-dimensional filter that produces the same results as applying TIFF0007742041000004.tif1030 vertically and then horizontally. The same pre-filter F can be obtained by applying a one-dimensional filter in the horizontal direction and then applying it in the vertical direction. The pre-filter F is obtained by multiplying the left and right matrices: It can also be expressed as TIFF0007742041000005.tif24127. The pre-filter F may be modified to perform equivalent calculations using the Hadamard transform. The prefilter F is the Hadamard matrix TIFF0007742041000006.tif1938 and the inverse Hadamard matrix Using TIFF0007742041000007.tif1936, It can be transformed into TIFF0007742041000008.tif19110. Therefore, pre-filtering for the four pixels included in the small block or filter block FB is performed as follows: (1) Hadamard transform (2) The high-frequency components α and β are multiplied by λ, and the high-frequency component γ is multiplied by λ. 2 double (3) Inverse Hadamard transform This can be done by an equivalent calculation using the Hadamard transform: Therefore, pre-filtering can be performed using logic circuits and program modules that can execute the Hadamard transform and inverse Hadamard transform, without the need for separate logic circuits and program modules for pre-filtering. The same applies to the post-filter described below. The Hadamard transform will now be explained. For the Hadamard transform of 2x2 pixels, the Hadamard matrix is Hadamard transform applied to TIFF0007742041000009.tif2037 TIFF0007742041000010.tif18150 gives μ, α, β, and γ. μ is the DC component (low frequency component) of the small block. The DC component of a pixel block is the average value of the pixel values ​​(luminance values) of all pixels contained in the pixel block. α, β, and γ are the horizontal AC component, vertical AC component, and diagonal AC component of the small block, respectively. The encoding process of the target image by the encoder 10 has been described above.

[0032] As a result of the encoding process, encoded data such as that shown in Figure 8(D) is generated. The target image has, for each block 50, a region that is intra-coded, a region that is inter-coded, and a region that is recursively coded by region division. For example, the upper left block 50 of the target image shown in Figure 8(A) is divided into four regions, and whether intra prediction, inter prediction, or hierarchical prediction will be used is determined for each region, so that there are portions of block 50 for which different prediction methods are selected (the prediction mode for each layer is not indicated, and block 50 as a whole is described as hierarchical inter coding). Since an appropriate coding method is applied to the coded data for each region, coding efficiency can be significantly improved compared to when inter-coding is performed in units of blocks 50 or larger regions.

[0033] Next, we will explain the decoding process of the coded data by the decoder 30. The decoding process is not simply the reverse of the process by the encoder 10. First, the second mode determination unit 43 of the decoder 30 determines, for each block 50 in the encoded data shown in Figure 8 (D), whether the prediction mode of the encoding mode is based on intra prediction, inter prediction, or recursive prediction that divides areas. If the prediction mode of the selected block is based on intra prediction, the decoder 30 decodes the selected block 50 using intra prediction; if the prediction mode of the selected block is based on inter prediction, the decoder 30 decodes the selected block 50 using inter prediction; if the prediction mode of the selected block 50 is based on recursive prediction that divides regions, the decoder 30 divides the block 50 into four regions that are half the size vertically and horizontally and decodes them recursively.

[0034] The second mode determination unit 43 of the decoder 30 determines whether a pre-filter has been applied to the selected block 50 or not. The pre-filter application determination can be made in a number of ways, including including information about whether or not a pre-filter is applied in the data, or determining that a pre-filter is applied if the block 50 is divided into four, and determining that a pre-filter is not applied if the block 50 is uniformly coded using intra-prediction or inter-prediction. By using the latter method, it is possible to standardize the code that expresses the mode in the pre-filter application conditions and prediction method, thereby reducing the amount of coding. The decoder 30 then performs decoding according to the prediction mode, outputs the decoded image of the block to which the prefilter is not applied as is, and performs decoding processing on the block to which the prefilter is applied after prefiltering the reference block, and then performs postfiltering. The decoded image contains boundary areas where pixel values ​​(hereafter referred to as "image") of blocks that underwent high-frequency emphasis (pre-filtering) during encoding are mixed with pixel values ​​(hereafter referred to as "true value") of blocks that did not undergo high-frequency emphasis, so a post-filter according to the pattern must be applied to the boundary areas.

[0035] FIG. 13 is a diagram illustrating a boundary pattern of post-filtering. The post filter is The file is TIFF0007742041000011.tif23150. FIG. 13(a) shows a 2×2 pixel block in which one pixel is the image, where pixel d is the image. At this time, the post-filter finds the true value d from {a, b, c, d'} (d' represents the image of d). The above prefilter F is calculated using u and the submatrix * that is not used in the calculation. If you put TIFF0007742041000012.tif1140, TIFF0007742041000013.tif1454 is obtained. It turns out that an inverse filter that obtains d from d', a, b, and c exists when (1+λ)≠0. The same applies when the image is not d but one of {a, b, c}.

[0036] FIG. 13(b) shows a 2×2 pixel block in which two pixels form an image, where pixels c and d form the image. At this time, the post filter finds the true values ​​{c,d} from {a,b,c',d'}. The above pre-filter F Place it as TIFF0007742041000014.tif1125, Solving for TIFF0007742041000015.tif7150, we get The resulting file is TIFF0007742041000016.tif1149. The inverse filter to obtain c and d from c', d', a, and b is detA=λ(1+λ). 2 It is known to exist when ≠ 0. The same is true if two other elements of {a, b, c, d} other than c and d are images.

[0037] FIG. 13(c) shows a case where three pixels form an image, and in this case, pixels b, c, and d form the image. At this time, the post filter finds the true value {b,c,d} from {a,b',c',d'}. Transforming the above prefilter F Place it as TIFF0007742041000017.tif1236, Solving for TIFF0007742041000018.tif13150, we get The resulting file is TIFF0007742041000019.tif1748. The inverse filter to obtain b, c, and d from b', c', d', and a is detC=16λ 2 +32λ 3 +16λ 4 It can be seen that it exists when ≠ 0. The same is true if three other elements of {a, b, c, d} other than b, c, and d are images. Therefore, it is clear that there exists a post-filter (inverse filter) for any boundary region where images and true values ​​are mixed. An appropriate inverse filter can be applied depending on the pattern of images and true values ​​in the boundary region.

[0038] Also, as shown in FIG. 13(d), when the pre-filter is applied to all 2×2 pixel blocks and all pixels are images, the post-filter (F -1 ) is applied to find the true value of each pixel. In the above explanation, the pre-filtering and Hadamard transform performed by the encoder 10 were described as separate steps. However, this is not limiting, and the encoder 10 may be implemented using a coefficient sequence or wavelet that can perform equivalent calculations by combining these steps. The same applies to the inverse Hadamard transform and post-filtering by the decoder 30, and the decoder 30 may be implemented using a coefficient sequence or wavelet that can perform equivalent calculations by combining these.

[0039] <Experimental Results> The inventor conducted an experiment to verify the encoding efficiency by encoding the same image data in three cases: when no pre-filtering was performed, when pre-filtering was performed on the entire target image, and when using a method of performing pre-filtering only on divided areas as in this embodiment. In the experiment, for the method of this embodiment, a pre-filter with odd-numbered pixels as vertices and a pre-filter with even-numbered pixels as vertices were applied twice in total to the entire image plane of the target image in inter-coding, and a post-filter was applied to the blocks to which the pre-filters were applied after decoding.Similarly, for the reference image, a pre-filter with odd-numbered pixels as vertices and a pre-filter with even-numbered pixels as vertices were applied twice in total to the entire image plane, and a post-filter was applied to the blocks to which the pre-filters were applied after decoding. This is because, unlike still images, inter-coding involves motion compensation, and therefore the predicted block obtained from the reference image does not necessarily have even-numbered pixels as its vertices. On the other hand, in a block of the target image, a pre-filter with odd-numbered pixels as its vertices always emphasizes pixel values ​​on the block boundary, while a pre-filter with even-numbered pixels as its vertices emphasizes the interior of the block. Therefore, the pre-filter with even-numbered pixels as its vertices is completely canceled by adjusting the 2x2 pixel Hadamard transform coefficient values, and does not contribute to the reduction of quantization noise. In a predicted block, there is no difference in the effect of a pre-filter with odd-numbered pixels as its vertices and a pre-filter with even-numbered pixels as its vertices. In addition, among the experimental results shown in Figures 15 to 19, when filtering is performed on the entire image, pre-filtering is performed on the entire target image before dividing the target image into blocks during encoding, and post-filtering is performed on the entire restored image after decoding during decoding.

[0040] FIG. 14 shows image data (1) to (5) used in the experiment, and in each image data, the left side is the target image and the right side is the reference image. Data (1) (CG, 1024 x 768 pixels) is image data of a scene in which a subject is rapidly deforming and zooming out. Data (2) (natural image, 720×480 pixels) is image data of a scene in which the background moves to the right and multiple subjects move in independent directions. Data (3) (natural image, 320 × 224 pixels) is a scene in which the subject moves to the right against a fixed background and the fountain changes rapidly. Data (4) (CG, 688 x 464 pixels, (c) copyright 2008, Blender Foundation / www.bigbuckbunny.org) is image data of a scene in which a subject zooms in slowly. Data (5) (natural image, 320×240 pixels) is image data of a scene in which the subject and background change at a slow rate. In the experiments, the prediction residual is quantized after Hadamard transformation and then Huffman coded. The sign of the value is coded with a fixed length of 1 bit, and other parameters are coded with Huffman coding. Motion compensation is performed with quarter-pixel accuracy, and fractional pixels are interpolated using a bilinear filter. The value of λ in the transformation matrix F is 1.25.

[0041] 15 to 19 are graphs showing the measurement results. FIG. 15 shows the measurement results using data (1), FIG. 16 shows the measurement results using data (2), FIG. 17 shows the measurement results using data (3), FIG. 18 shows the measurement results using data (4), and FIG. 19 shows the measurement results using data (5). The vertical axis in Figures 15 to 19 is PSNR ([dB]), with a higher value indicating less distortion and better performance. The horizontal axis is the compression ratio, with a lower value indicating better performance, with less coding required to describe the data compared to the original data. In general, the closer the data is plotted to the upper left of the graph, the better the codec. It can be seen that in scenes in data (1) to (3) where the subject moves or deforms and the correlation is low, the coding efficiency is improved by the effect of filtering. In particular, according to the method of this embodiment, the PSNR is improved by about 0.8 [dB] in the high-image-quality region of data (1), and by about 0.5 [dB] in the high-image-quality region of data (2). At the same time, for data (4) and (5) that are highly correlated and do not involve movement or deformation of the subject, degradation is more effectively suppressed than in a method in which pre-filtering is performed on the entire image and then post-filtering is performed on the entire image during decoding. In other words, when predicting a low correlation between a target image and a reference image, filtering the entire image improves image quality and coding efficiency compared to not filtering, but when the correlation is high, not filtering improves image quality and coding efficiency. Therefore, filtering can produce good or bad results depending on the conditions of the image being processed. However, by filtering on a block-by-block basis as in the present invention, high-quality image processing with high coding efficiency can be achieved regardless of the conditions of the image being processed.

[0042] 20 is a flowchart illustrating the encoding process executed by the encoder. As described above, this embodiment employs an algorithm in which prediction is performed on a block that is not subjected to pre-filtering, and when the prediction residual is equal to or greater than a predetermined value in any prediction mode, pre-filtering is performed on the block of the target image to perform prediction. In step S101, the encoder 10 selects blocks into which the target image is divided. In step S102, the encoder 10 performs intra prediction on the selected block. In step S103, the encoder 10 determines whether or not there is a prediction residual in the intra prediction that is equal to or greater than a predetermined value. If it is determined that there is a prediction residual of intra prediction equal to or greater than the predetermined value (Yes in step S103), the encoder 10 performs motion compensation prediction (inter prediction) on the selected block in step S104. Then, in step S105, the encoder 10 determines whether or not there is a prediction residual in inter prediction that is equal to or greater than a predetermined value. If it is determined that there is a prediction residual equal to or greater than a predetermined value in inter prediction (Yes in step S105), the encoder 10 further divides the block in step S106. If it is determined that there is no residual error greater than or equal to a predetermined value in intra prediction (No in step S103), the encoder 10 determines the coding method for this block to be intra prediction in step S107, codes it using intra prediction, and proceeds to step S112. Also, if it is determined that there is no residual greater than a predetermined value in inter prediction (No in step S105), the encoder 10 determines the coding method for this block to be inter prediction in step S108, codes it using inter prediction, and proceeds to step S112.

[0043] In step S109, the encoder 10 performs pre-filtering on the blocks divided in step S106 and the reference block. In step S110, the encoder 10 performs hierarchical inter prediction on the blocks divided in step S106. In step S111, the encoder 10 encodes the prediction residuals of step S111, and proceeds to step S112. In step S112, the encoder 10 determines whether there are any unselected blocks. If it is determined that there is an unselected block (No in step S112), the encoder 10 returns to step S101 and repeats the processes of steps S101 to S111 for the unselected block. If it is determined that there are no unselected blocks and that the processing has been completed for all blocks (Yes in step S112), the encoder 10 ends the encoding process. Note that, in steps S107 and S108, the flow in Fig. 16 is based on the assumption that no residuals are generated when encoding using intra prediction or inter prediction. However, if residuals are generated, a processing step for encoding the residuals may be included after each of steps S107 and S108.

[0044] FIG. 21 is a flowchart illustrating the decoding process executed by the decoder. In step S201, the decoder 30 selects a block in the encoded data. In step S202, the decoder 30 determines whether the selected block is coded in its entirety (using the same coding method). If it is determined that the entire block has been coded (Yes in step S202), the decoder 30 determines in step S203 whether the coding method is intra prediction. If it is determined that the coding method is intra prediction (Yes in step S203), the decoder 30 decodes the block using intra prediction in step S204. If it is determined that the encoding method is not intra prediction (No in step S203), the decoder 30 decodes the block using inter prediction in step S205. In the decoding process in steps S204 and S205, residual information, if any, is also included in the decoding.

[0045] If it is determined in step S201 that the block selected has not been coded as a whole (No in step S202), the decoder 30 further divides the block (step S206). In step S207, the decoder 30 determines whether a prefilter has been applied to the block, and if a prefilter has been applied (Yes in step S207), performs prefiltering on the reference block (step S208). Thereafter, the block is decoded using hierarchical inter prediction (step S209), and if there is a residual, the residual is decoded (step S210), and the decoded block is post-filtered (step S211). On the other hand, if a pre-filter is not applied to the divided block (No in step S207), the block is decoded using hierarchical inter prediction (step S212), and if there is a residual, the residual is decoded (step S213). After the process of any one of steps S204, S205, S211, and S213, the decoder 30 determines in step S214 whether or not there are any unselected blocks. If it is determined that there is an unselected block (No in step S214), the decoder 30 returns the process to step S201 and performs the processes of steps S201 to S213 on the next block. If it is determined that there are no unselected blocks (Yes in step S214), the decoder 30 integrates the decoded blocks to restore the target image (step S215), and completes the decoding process. If an algorithm is used that determines that a pre-filter has been applied when a block is divided, the process of step S207 may be omitted (steps S212 and S213 may also be omitted).

[0046] FIG. 22 is a block diagram illustrating an embodiment of a computer system. The configuration of the computer device 100 will be described with reference to FIG. The computer device 100 is, for example, an image processing device that processes various types of information. The computer device 100 includes a control circuit 101, a storage device 102, a reading / writing device 103, a recording medium 104, a communication interface 105, an input / output interface 106, an input device 107, and a display device 108. The communication interface 105 is connected to a network 200. The components are connected by a bus 110. The image processing device 10 and the image processing device 30 can be configured by appropriately selecting some or all of the components described in the computer device 100. The control circuit 101 controls the entire computer device 100. The control circuit 101 is a processor such as a central processing unit (CPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or a programmable logic device (PLD). The control circuit 101 functions as, for example, the control unit 11 in FIG. 2 or the control unit 31 in FIG. 3. The storage device 102 stores various data. The storage device 102 is, for example, a memory such as a read-only memory (ROM) or a random access memory (RAM), a hard disk (HD), or a solid state drive (SSD). The storage device 102 may store an image processing program that causes the control circuit 101 to function as the control unit 11 in FIG. 2 or 6 or the control unit 31 in FIG. 3 or 7. The storage device 102 functions, for example, as the storage unit 12 in FIG. 2 or 6 or the storage unit 32 in FIG. 3 or 7. When performing image processing, the image processing device 10 and the image processing device 30 read out the program stored in the storage device 102 into the RAM.

[0047] In the first embodiment, the image processing device 10 executes a program read into the RAM in the control circuit 101, thereby performing processing including one or more of reception processing, prediction processing, first and second prefilter processing, encoding processing, and output processing. The image processing device 30 executes the program read into the RAM in the control circuit 101, thereby performing processing including one or more of reception processing, prediction processing, third pre-filter processing, post-filter processing, decoding processing, and output processing. In the second embodiment, the image processing device 10 executes a program read into the RAM in the control circuit 101, thereby performing processing including one or more of reception processing, division processing, first and second prediction processing, first mode determination processing, first and second prefilter processing, encoding processing, and output processing. The image processing device 30 executes the program read into the RAM in the control circuit 101, thereby performing processing including one or more of reception processing, combining processing, third and fourth prediction processing, second mode determination processing, third pre-filter processing, post-filter processing, decoding processing, and output processing. The program may be stored in a storage device of a server on the network 200 as long as the control circuit 101 can access it via the communication interface 105. The reading / writing device 103 is controlled by the control circuit 101 and reads / writes data from / to a removable recording medium 104 . The recording medium 104 stores various data. For example, the recording medium 104 stores an image processing program. The recording medium 104 is, for example, a non-volatile memory (non-transitory recording medium) such as a Secure Digital (SD) memory card, a Floppy Disk (FD), a Compact Disc (CD), a Digital Versatile Disk (DVD), a Blu-ray (registered trademark) Disk (BD), or a flash memory. The communication interface 105 communicably connects the computer device 100 to other devices via the network 200. The communication interface 105 functions as, for example, the communication unit 13 in FIG. 2 and the communication unit 33 in FIG.

[0048] The input / output interface 106 is an interface that is detachably connected to, for example, various input devices. Examples of the input devices 107 that are connected to the input / output interface 106 include a keyboard and a mouse. The input / output interface 106 communicatively connects the various connected input devices to the computer apparatus 100. The input / output interface 106 outputs signals input from the various connected input devices to the control circuit 101 via the bus 110. The input / output interface 106 also outputs signals output from the control circuit 101 to the input / output devices via the bus 110. The input / output interface 106 functions as, for example, the input unit 14 in FIG. 2 or the input unit 34 in FIG. 3. Display device 108 displays various types of information. Display device 108 is, for example, a CRT (Cathode Ray Tube), LCD (Liquid Crystal Display), PDP (Plasma Display Panel), or OELD (Organic Electroluminescence Display). Network 200 is, for example, a LAN, wireless communication, a P2P network, or the Internet, and communicatively connects computer device 100 to other devices. It should be noted that the present embodiment is not limited to the above-described embodiment, and various configurations or embodiments can be adopted within the scope of the present embodiment without departing from the gist of the present embodiment. [Explanation of symbols]

[0049] 1 Image processing system, 10 Image processing device (encoder), 30 Image processing device (decoder), 100 Computer device, 101 Control circuit, 102 Storage device, 103 Reading device, 104 Recording medium, 105 Communication interface, 106 Input / output interface, 107 Input device, 108 Display device, 110 Bus, 200 Network

Claims

1. A coding device for coding moving images, comprising: a division unit that divides an image to be encoded into blocks; a first mode determination unit that determines a prediction mode and whether or not pre-filtering is performed; a first pre-filter unit that performs pre-filtering on a target image to generate a pre-filtered target image block; a second pre-filter unit that performs pre-filtering on the reference image to generate a pre-filtered reference image block; a first prediction unit for predicting pixel values ​​of a block of a target image that is not pre-filtered; a second prediction unit that predicts pixel values ​​of the pre-filtered target image; an encoding unit that encodes a target image and generates encoded data; Equipped with The encoding device, characterized in that the second prediction unit performs any one of intra prediction, inter prediction, and hierarchical inter prediction using the output of the first pre-filter unit, or the outputs of the first pre-filter unit and the second pre-filter unit.

2. In the encoding device according to claim 1, The first pre-filter section and the second pre-filter section are pre-filtering based on even-numbered pixels of the target image block and the reference image block; pre-filtering based on odd-numbered pixels of the target image block and the reference image block; The encoding device is characterized by performing pre-filtering twice.

3. A second mode determination unit that determines a prediction mode of encoded data and whether or not pre-filtering is performed; a third prediction unit that predicts pixel values ​​of blocks of a target image that are not pre-filtered during encoding; a fourth prediction unit that predicts pixel values ​​of a block of a target image that has been prefiltered during encoding; a third pre-filter unit that performs pre-filtering on the reference image to generate a pre-filtered reference image block; a decoding unit that reconstructs a block by adding a predetermined predicted value based on a prediction mode and a coded residual to the coded data; a postfilter unit that performs predetermined post-filtering on the decoded block to generate a post-filtered block; a combining unit for combining the post-filtered blocks to obtain a restored image; 2. A decoding device according to claim 1, wherein when the fourth prediction unit performs inter-prediction, the fourth prediction unit performs prediction using an output from the third pre-filter unit.

4. In the decoding device according to claim 3, the post-filter unit uses different post-filters at boundary portions where pre-filtering and non-pre-filtering are mixed; A decoding device characterized by:

5. A processor-implemented encoding method for encoding an image, comprising: Dividing an image to be encoded into blocks; determining a prediction mode; performing pre-filtering on the block of the target image to generate a block of the pre-filtered target image; performing pre-filtering on the block of the reference image to generate a pre-filtered block of the reference image; Inter-predicting pixel values ​​of blocks of the pre-filtered target image by referring to blocks of the pre-filtered reference image; generating coded data by encoding a predetermined predicted value based on the prediction mode and a pixel value of the block or a residual of the pixel value of the block of the pre-filtered target image and the prediction mode; 10. A coding method comprising:

6. An image processing program that causes a processor to execute an encoding method for encoding an image, Dividing an image to be encoded into blocks; determining a prediction mode; performing pre-filtering on the block of the target image to generate a block of the pre-filtered target image; pre-filtering the block of the reference image to generate a pre-filtered block of the reference image; Inter-predicting pixel values ​​of blocks of the pre-filtered target image by referring to blocks of the pre-filtered reference image; generating coded data by encoding a predetermined predicted value based on the prediction mode and a pixel value of the block or a residual of the pixel value of the block of the pre-filtered target image and the prediction mode; 10. An encoding program comprising:

7. A decoding method for decoding an encoded image, the method being performed by a processor, comprising: determining a prediction mode of the encoded data; performing pre-filtering on the reference image to generate a block of pre-filtered reference image; Inter-predicting pixel values ​​of blocks of the pre-filtered target image by referring to blocks of the pre-filtered reference image; Reconstructing the block by adding a predetermined prediction value based on a prediction mode and a coded residual to the coded data; generating a post-filtered block by performing predetermined post-filtering on the restored block; reconstructing the image by combining the post-filtered blocks; A decoding method comprising:

8. An image processing program that causes a processor to execute a decoding method for decoding an encoded image, comprising: determining a prediction mode of the encoded data; performing pre-filtering on the reference image to generate a block of pre-filtered reference image; Inter-predicting pixel values ​​of blocks of the pre-filtered target image by referring to blocks of the pre-filtered reference image; Reconstructing the block by adding a predetermined prediction value based on a prediction mode and a coded residual to the coded data; generating a post-filtered block by performing predetermined post-filtering on the restored block; reconstructing the image by combining the post-filtered blocks; A decoding program characterized by:

Citation Information

Patent Citations

  • Image encoder and image encoding method amd medium recording program describing the method

    JP2002077909A

  • Image coder and image coding method

    JP2002152758A

  • Method for encoding moving image

    JP2002247576A

  • Video signal coding instrument, video signal coding method, mobile terminal instrument, and video signal coding program

    JP2005168053A

  • Method and apparatus for adaptive combined pre-processing and post-processing filters for video coding and decoding.

    JP2013516834A