Image decoding device and image decoding method

By applying low-pass filtering and block division with DST or DCT selection, the method enhances video encoding efficiency by reducing signal errors and residual components in prediction residual signals, addressing inefficiencies in existing video encoding technologies.

JP7837442B2Active Publication Date: 2026-03-30NIPPON HOSO KYOKAI
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2026-03-30

AI Technical Summary

Technical Problem

Existing video encoding methods, such as MPEG-2, AVC/H.264, MPEG-H, and HEVC/H.265, do not effectively utilize the high correlation between adjacent pixel signals in intra-prediction, leading to inefficiencies in encoding and decoding processes, particularly at the edges of prediction regions.

Method used

An image encoding and decoding device that applies low-pass filtering to prediction signals using adjacent decoded signals, followed by block division and selection of either Discrete Sine Transform (DST) or Discrete Cosine Transform (DCT) based on the direction of the filter, to reduce prediction residual signals and improve encoding efficiency.

Benefits of technology

The proposed method reduces signal errors and improves encoding efficiency by minimizing the residual component in prediction residual signals, allowing for high-efficiency video encoding and decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007837442000001
    Figure 0007837442000001
  • Figure 0007837442000002
    Figure 0007837442000002
  • Figure 0007837442000003
    Figure 0007837442000003
Patent Text Reader

Abstract

To provide an image encoding device, an image decoding device, and a program for improving encoding efficiency.SOLUTION: An image decoding device that decodes signals that have been coded by dividing frames that constitute a moving image into blocks includes a prediction unit that generates block-based predicted images including predicted signals by performing signal prediction on each pixel signal on a block-by-block basis, an acquisition unit that acquires a control identification signal that indicates the type of selection control selected by the coding side, a filtering unit that generates a new block of the predicted image by filtering the predicted signal using decoded adjacent signals adjacent to the block of the predicted image on the basis of the control identification signal, an inverse quantization unit that performs corresponding inverse quantization processing on quantized transform coefficients to restore transform coefficient signals, and an inverse orthogonal transform unit that reconstructs block-by-block predicted residual signals by performing inverse transform processing determined on the basis of the control identification signal on the transform coefficient signals.SELECTED DRAWING: Figure 17
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image encoding apparatus, an image decoding apparatus, and programs thereof applicable to video encoding methods such as MPEG-2, AVC / H.264, MPEG-H, and HEVC / H.265.

Background Art

[0002] It is generally known that image signals, regardless of still images or moving images, have a high signal correlation between adjacent pixels. Utilizing this property, for example, in video encoding methods such as MPEG-2, AVC / H.264, MPEG-H, and HEVC / H.265, two prediction modes, intra prediction and inter prediction, are provided. Intra prediction is a technique for performing signal prediction using only the signals within an encoding frame. For the pixel signals (original signals) within the encoding target block of the original image, DC prediction, Planer prediction, or directional prediction is performed using the pixel signals of the encoded and decoded blocks adjacent to the left or upper side of the encoding target block, and a block of a predicted image composed of pixel signals (predicted signals) predicted by extrapolation is generated. Thus, the pixel signals within the encoding target block are efficiently predicted by using the pixel signals of the decoded blocks adjacent to the encoding target block.

[0003] For example, referring to FIG. 29, the nature of the prediction residual signal by intra prediction will be described using the horizontal prediction in intra prediction as an example. In the example shown in FIG. 29, for the original signal within the encoding target block Bko of the original image composed of pixel signals p with a block size of horizontal 4 pixels × vertical 4 pixels (hereinafter simply referred to as "4×4", and the same applies to other block sizes) (see FIG. 29(a)), prediction is performed using the pixel signals of the decoded block adjacent to the left as the reference signal Sr for horizontal prediction, and a block Bkp of a predicted image composed of predicted signals corresponding to the pixel signals of the encoding target block is generated (see FIG. 29(b)).

[0004] Then, the predicted residual signal block Bkd can be obtained from the difference between the original signal in the encoded block Bko of the original image and the predicted signal in block Bkp of the predicted image (see Figure 29(c)). This intra-predicted residual signal has the property that the prediction efficiency is higher for pixels adjacent to the reference signal, so the residual component becomes smaller (signal intensity is lower), and in other words, the prediction efficiency decreases as you move away from the reference signal, so the residual component becomes larger (signal intensity is higher) (see Figure 29(d)).

[0005] In the example shown in Figure 29, prediction is performed using horizontal inter-pixel correlation, but the signal of the encoded and decoded block also exists above the block being encoded. In the example shown in Figure 29, horizontal intra-prediction efficiently uses horizontal correlation to make predictions and generate prediction residual signals, but does not utilize vertical correlation.

[0006] Therefore, in the currently defined H.265, a filter is applied to only the uppermost predicted signal within block Bkp of the predicted image, taking advantage of the high correlation with the pixel signals in the block above the block being encoded. This allows for the use of not only the horizontal correlation from horizontal prediction but also the vertical correlation, further reducing the average residual component in block Bkd of the predicted residual signal and improving encoding efficiency. Similarly, in the case of vertical intra-prediction, a filter is applied using horizontal adjacent pixels that were not used during the prediction, and in DC prediction, which does not use horizontal and vertical adjacent pixels, a filter is applied using horizontal and vertical adjacent pixels, resulting in a significant improvement in encoding efficiency compared to previous video encoding methods such as MPEG-2 and H.264.

[0007] Furthermore, in interpretation, motion is predicted from temporally adjacent encoded and decoded reference frames to calculate motion vectors, which are then used to generate a predicted image. The difference between the predicted image and the original image is generated as a predicted residual signal. However, while conventional techniques actively utilize the high correlation between signals in adjacent blocks in intrapretation, interpretation does not utilize this property of high correlation between signals in adjacent blocks.

[0008] The blocks of predicted residual signals generated through intra-prediction and inter-prediction are then subjected to orthogonal transformation and quantization processes and encoded.

[0009] Currently, the orthogonal transformation processing for intra-prediction in H.265 primarily uses the Discrete Cosine Transform (DCT), but a mode is available that applies the Discrete Sine Transform (DST) orthogonal transformation processing to the prediction residual signals of intra-prediction for certain block sizes. However, if the block size of the block to be encoded is large (larger than 4x4), the DST will be applied based on the signal features of the blocks to its left and above. Due to reduced benefits and implementation cost issues, the application of DST is not permitted by the standard. Furthermore, in the orthogonal transform processing of interpretation as currently defined in H.265, DCT is applied regardless of block size.

[0010] A characteristic of motion compensation processing in interpretation is that the prediction error is statistically large at the endpoints of the prediction region involved in its encoding (i.e., the prediction residual signal for pixel signals located on the outer periphery of the predicted image block) (see, for example, Non-Patent Document 1). Therefore, when this prediction region is further divided into blocks to form a transformation region, there exists a transformation region whose features match those of the base DST used when encoding the prediction residual signal. A technique for applying DST using this has been disclosed (see, for example, Patent Document 1). [Prior art documents] [Patent Documents]

[0011] [Patent Document 1] Japanese Patent Publication No. 2014-36278 [Non-patent literature]

[0012] [Non-Patent Document 1] Zheng Wen-tao, et al., "Characteristic Analysis of Motion-Compensated Interframe Difference Signals Based on a Statistical Motion Distribution Model," D-II, Vol.J84-D-II, No.9, pp.2001-2010, September 1, 2001. [Overview of the project] [Problems that the invention aims to solve]

[0013] As mentioned above, Non-Patent Document 1 shows that a characteristic of motion compensation processing in interpretation is that the amount of prediction error is statistically large at the endpoints of the prediction region related to the encoding (i.e., the prediction residual signal for pixel signals located on the outer edge of the block of the prediction image), and based on this, Patent Document 1 discloses a technique for applying DST.

[0014] However, since this is statistically proven, applying DST to statistically outlier images (for example, when there are features such as strong edges near the outer edges of the predicted image blocks) can actually decrease encoding efficiency.

[0015] The object of the present invention is to provide an image coding device, an image decoding device, and programs thereof that improve coding efficiency. [Means for solving the problem]

[0016] The image decoding device according to the present invention is an image decoding device that decodes a signal encoded by dividing a frame constituting a moving image into blocks, and is characterized by comprising: a prediction unit that generates a block-unit predicted image including a predicted signal by performing signal prediction for each pixel signal in block units; an acquisition unit that acquires a control identification signal indicating the type of selection control selected by the encoding side; a filter processing unit that generates a new block of predicted image by applying a filter process to the predicted signal using decoded adjacent signals adjacent to the block of the predicted image based on the control identification signal; an inverse quantization unit that restores a conversion coefficient signal by applying a corresponding inverse quantization process to the quantized conversion coefficients; and an inverse transformation unit that reconstructs a block-unit predicted residual signal by applying an inverse transformation process determined based on the control identification signal to the conversion coefficient signal. In the image encoding process in one embodiment, a low-pass filter is applied to the prediction image used for interpretation, similar to intraprediction, using the encoded and decoded adjacent signals. This reduces the prediction residual signal for the leftmost and uppermost regions of the prediction image used for interpretation. The prediction residual signal is then divided into blocks, and either DST or DCT is determined without a flag depending on the direction of the low-pass filter, and the corresponding orthogonal transformation is executed. In the image decoding process in one embodiment, either DST or DCT is determined without a flag depending on the direction of the low-pass filter, and the original image is restored.

[0017] In other words, an image encoding device according to one embodiment is an image encoding device that divides and encodes a frame-by-frame original image constituting a moving image into blocks, and is characterized by comprising: adjacent pixel non-reference prediction means that generates a block-unit predicted image consisting of predicted signals by predetermined adjacent pixel non-reference prediction that performs signal prediction for each pixel signal of the block-unit original image without using decoded adjacent signals; filtering means that generates a new block of predicted image by applying a low-pass filter to the predicted signals located at the boundary of the block of the predicted image using the decoded adjacent signals adjacent to the block of the predicted image; predicted residual signal generation means that calculates the error of each predicted signal in the block of the predicted image after the low-pass filter processing for each pixel signal of the block to be encoded of the original image, and generates a block-unit predicted residual signal; and orthogonal transformation means having block division means that divides the block-unit predicted residual signal into a predetermined block shape, and orthogonal transformation selection application means that selects and applies a plurality of orthogonal transformation processes to each divided block of the divided predicted residual signal according to the position to which the low-pass filter processing was applied.

[0018] Furthermore, in an image encoding apparatus of one embodiment, the predetermined adjacent pixel non-reference prediction is characterized by including any of inter prediction, intrablock copy prediction, and cross-component signal prediction.

[0019] Furthermore, in an image encoding apparatus of one embodiment, the filtering means is characterized in that it applies the low-pass filter processing to predict the predicted signals adjacent to the decoded adjacent signals among the predicted signals within the block of the predicted image generated by the adjacent pixel non-reference prediction means, so as to be smoothed.

[0020] Furthermore, in an image encoding apparatus of one embodiment, the block division means in the orthogonal transformation means is characterized in that, when dividing blocks of the predicted residual signal corresponding to blocks of the predicted image with a block size larger than twice the size obtained by multiplying the block size by a predetermined number of times in length and width, the division means divides the division block at a position adjacent to the decoded adjacent signal so that it becomes the predetermined block size.

[0021] Furthermore, in an image encoding apparatus of one embodiment, the block division means in the orthogonal transformation means is characterized in that the divided blocks including the pixel positions to which the low-pass filter processing is applied are of a predetermined block size.

[0022] Furthermore, in an image encoding apparatus of one embodiment, the orthogonal transformation selection application means in the orthogonal transformation means is characterized in that, among the divided blocks of the divided predicted residual signal, when the pixel to which the low-pass filter processing is applied is at the uppermost position and the direction of application of the low-pass filter processing is vertical, a first orthogonal transformation process of vertical DST and horizontal DCT is applied to the divided block of the predicted residual signal located at the uppermost position; when the pixel to which the low-pass filter processing is applied is at the leftmost position and the direction of application of the low-pass filter processing is horizontal, a second orthogonal transformation process of horizontal DST and vertical DCT is applied to the divided block of the predicted residual signal located at the leftmost position; when the pixel to which the low-pass filter processing is applied is at the uppermost position and at the leftmost position, and the direction of application of the low-pass filter processing is the direction of application of the angular filter processing, a third orthogonal transformation process of vertical and horizontal DST is applied; and for divided blocks of the predicted residual signal where there are no pixels to which the low-pass filter processing is applied, a fourth orthogonal transformation process of vertical and horizontal DCT is applied.

[0023] Also, in the image encoding device according to an embodiment, the filter processing means acquires a local decoded image for each divided block based on the block dividing means, and each time the local decoded image is acquired, replaces the prediction signal in the block of the prediction image generated by the non-reference prediction of adjacent pixels, and for each block size of the divided block, among the prediction signals in the block of the replaced prediction image, uses the decoded adjacent signal adjacent to the prediction image and the local decoded image for each divided block to perform a low-pass filter process according to each block size.

[0024] Furthermore, an image decoding device according to an embodiment is an image decoding device that decodes a signal obtained by block-dividing frames constituting a moving image, and includes non-reference prediction means for adjacent pixels that generates a prediction image in units of blocks composed of prediction signals by performing signal prediction without using a decoded adjacent signal for each pixel signal in units of blocks, filter processing means that generates a new block of the prediction image by performing a low-pass filter process on the prediction signal located at the boundary of the block of the prediction image using the decoded adjacent signal adjacent to the block of the prediction image, orthogonal transformation selection application means that selects and applies a plurality of types of inverse orthogonal transformation processes corresponding to the position where the low-pass filter process is applied to the conversion coefficient that was block-divided by the image encoding side for the restored conversion coefficient in units of blocks to generate a divided block of the prediction residual signal, and inverse orthogonal transformation means including block reconstruction means that reconstructs a block corresponding to the block size of the prediction image based on the divided block of the prediction residual signal.

[0025] Also, in the image decoding apparatus according to one embodiment, the filter processing means acquires a local decoded image for each divided block based on the block dividing means, and each time the local decoded image is acquired, the prediction signal within the block of the prediction image generated by the non-reference prediction of adjacent pixels is replaced, and within the block of the replaced prediction image in units of the block size of the divided block, among the prediction signals within the block of the replaced prediction image, the decoded adjacent signal adjacent to the prediction image and the local decoded image for each divided block are used to perform a low-pass filter process corresponding to each block size.

[0026] Furthermore, a program according to one embodiment causes a computer to function as the image encoding apparatus or the image decoding apparatus according to one embodiment.

Advantages of the Invention

[0027] According to one embodiment, since the signal error generated between the prediction signal of the prediction image and the adjacent decoded block signal is reduced and the encoding efficiency is improved, an image encoding apparatus and an image decoding apparatus of a video encoding method with high encoding efficiency can be realized. That is, on the encoding side, the residual component in the prediction residual signal can be made smaller, and the encoding efficiency can be improved by reducing the amount of information to be encoded and transmitted. Also, on the decoding side, decoding can be performed even with the reduced amount of information.

Brief Description of the Drawings

[0028] [Figure 1] (a) is a block diagram around a filter processing unit related to a prediction image in the image encoding apparatus according to the first embodiment of the present invention, and (b) is an explanatory diagram showing an example of the filter processing of the prediction image. [Figure 2] (a), (b), and (c) are explanatory diagrams respectively illustrating non-reference prediction of adjacent pixels according to the present invention. [Figure 3] It is a block diagram showing an example of an image encoding apparatus according to the first embodiment of the present invention. [Figure 4]This is a flowchart relating to the filtering process of predicted images in the image encoding device of the first embodiment according to the present invention. [Figure 5] This is an explanatory diagram relating to the filtering process of a predicted image in an image encoding device according to the first embodiment of the present invention. [Figure 6] This is a block diagram of the area surrounding the filter processing unit related to the predicted image in the image decoding device of the first embodiment of the present invention. [Figure 7] This is a block diagram showing one embodiment of an image decoding device according to the first embodiment of the present invention. [Figure 8] This is a block diagram of a filter processing unit in an image encoding device or image decoding device according to a second embodiment of the present invention. [Figure 9] This is a flowchart relating to the filtering process of predicted images in an image encoding device or image decoding device according to a second embodiment of the present invention. [Figure 10] (a) and (b) are explanatory diagrams relating to the filtering process of a predicted image in an image encoding device or image decoding device of a second embodiment according to the present invention. [Figure 11] (a) is a block diagram of the filter processing unit and the area around the orthogonal transform selection control unit in the image encoding device of the third embodiment of the present invention, and (b) and (c) are diagrams showing examples of base waveforms of the orthogonal transform. [Figure 12] This is a block diagram showing the first embodiment of the image encoding device according to the third embodiment of the present invention. [Figure 13] This is a flowchart relating to the orthogonal transform selection control process of the first embodiment in the image encoding device of the third embodiment according to the present invention. [Figure 14] This is a block diagram showing a second embodiment of the image encoding device according to the third embodiment of the present invention. [Figure 15] This is a flowchart relating to the orthogonal transform selection control process of the second embodiment in the image encoding device of the third embodiment of the present invention. [Figure 16] This is a block diagram of the filter processing unit and the area around the inverse orthogonal transform selection control unit in the image decoding device of the third embodiment of the present invention. [Figure 17] This is a block diagram showing one embodiment of an image decoding device according to the third embodiment of the present invention. [Figure 18] This is a flowchart relating to the inverse orthogonal transform selection control process in one embodiment of the image decoding device according to the third embodiment of the present invention. [Figure 19] (a) is a block diagram of the area around the filter processing unit for the predicted image in the image encoding device of the fourth embodiment of the present invention, and (b) is an explanatory diagram showing an example of the filter processing of the predicted image. [Figure 20] (a), (b), and (c) are explanatory diagrams illustrating the application of orthogonal transformations to blocks after block division of a predicted image according to the present invention. [Figure 21] This is a block diagram showing one embodiment of an image encoding device according to the fourth embodiment of the present invention. [Figure 22] This is a flowchart relating to the filtering process of predicted images in the image encoding device of the fourth embodiment according to the present invention. [Figure 23] This is a flowchart relating to the orthogonal transformation process in the image encoding device of the fourth embodiment according to the present invention. [Figure 24] This is a block diagram of the area surrounding the filter processing unit related to the predicted image in the image decoding device of the fourth embodiment of the present invention. [Figure 25] This is a block diagram showing one embodiment of an image decoding device according to the fourth embodiment of the present invention. [Figure 26] (a) and (b) are block diagrams of the area around the filter processing unit in the image encoding device and image decoding device of the fifth embodiment according to the present invention, respectively. [Figure 27] This is a flowchart relating to the filtering process of predicted images in an image encoding device or image decoding device according to the fifth embodiment of the present invention. [Figure 28] This is an explanatory diagram relating to the filtering process of a predicted image in an image encoding device or image decoding device according to the fifth embodiment of the present invention. [Figure 29](a), (b), (c), and (d) are explanatory diagrams regarding intra-prediction in conventional techniques. [Modes for carrying out the invention]

[0029] The image encoding apparatus and image decoding apparatus of each embodiment of the present invention will be described in order below.

[0030] [First Embodiment] (Image encoding device) First, the filter processing unit 12 related to the predicted image, which is a main component of the image encoding device 1 according to the first embodiment of the present invention, will be described with reference to Figures 1 and 2, and a specific typical example of the image encoding device 1 will be described with reference to Figures 3 to 5.

[0031] Figure 1(a) is a block diagram of the area around the filter processing unit 12 related to the predicted image in the image encoding device 1 of the first embodiment according to the present invention, and Figure 1(b) is an explanatory diagram showing an example of filtering the predicted image by the filter processing unit 12.

[0032] As shown in Figure 1(a), the image encoding apparatus 1 of the first embodiment according to the present invention is configured to include an adjacent pixel non-reference prediction unit 11, a filter processing unit 12, a prediction residual signal generation unit 13, and an orthogonal transformation unit 14.

[0033] The adjacent pixel non-reference prediction unit 11 is a functional unit that performs signal prediction processing to generate a predicted image for the encoding target region of the original image without using adjacent decoded signals (decoded adjacent signals). In this specification, such signal prediction is referred to as "adjacent pixel non-reference prediction," and adjacent pixel non-reference prediction includes, for example, inter-prediction which performs motion compensation prediction between frames, intra-block copy prediction which generates a predicted image by copying different decoded partial images of the same frame, and signal prediction which generates a predicted image by utilizing the correlation between component signals, namely the luminance signal and the chrominance signal, at corresponding block positions in a certain frame (referred to as "cross-component signal prediction" in this specification).

[0034] For example, in interpretation, as shown in Figure 2(a), in a GOP structure with multiple frames F1 to F6, if frame F1 is an I-frame (intra-frame), then a predicted image may be generated by referencing past I-frames and P-frames, such as a P-frame (predictive inter-frame), or by referencing multiple frames such as past and future frames, such as a B-frame (dual-predictive inter-frame).

[0035] Furthermore, in intrablock copy prediction, as shown in Figure 2(b), a predicted image Bkp is generated by referencing and copying different decoded partial images Bkr from the same frame F.

[0036] Furthermore, in cross-component signal prediction, as shown in Figure 2(c), the correlation between component signals is utilized, and a local decoded image Bkr of the luminance signal (e.g., Y signal) at a corresponding block position in a certain frame F is referenced, and a weighted addition synthesis process is applied to the predicted image of the chrominance signal (e.g., U / V signal) to generate a corrected predicted image Bkp of the chrominance signal (e.g., U / V signal). For further details regarding this cross-component signal prediction, please refer to Japanese Patent Application Publication No. 2014-158270.

[0037] These neighboring pixel non-reference predictions share a common characteristic: they predict the signal without utilizing the high correlation between neighboring pixels, even though they have a decoded region adjacent to the encoded region block of the original image.

[0038] In the first embodiment, the filter processing unit 12 is a functional unit that generates a new predicted image consisting of a new predicted signal by applying a low-pass filter to the predicted signal using the decoded signals (decoded adjacent signals) that are adjacent to the left and above the predicted image, respectively, from the pixel signals (predicted signals) within the block of the predicted image generated by the adjacent pixel non-reference prediction unit 11, and outputs this to the predicted residual signal generation unit 13.

[0039] For example, as shown in Figure 1(b), if the filter processing unit 12 assumes that block Bkp of the predicted image, which consists of 4x4 pixel signals p corresponding to block Bko of the encoding target region of the original image, is generated by the adjacent pixel non-reference prediction unit 11, the filter processing unit 12 uses the decoded adjacent signals So around the block boundary of block Bkp of the predicted image, which are adjacent to the left and above the predicted image, to apply low-pass filtering, such as horizontal and vertical smoothing filters, to the leftmost and uppermost pixel signals (predicted signals) of the predicted image as signals Sf of the filtering target region, respectively, thereby generating a new predicted image consisting of the predicted signals.

[0040] The prediction residual signal generation unit 13 calculates the error between each pixel signal (predicted signal) of block Bkp of the predicted image obtained from the filter processing unit 12 and each pixel signal (original signal) of block Bko of the encoding target region (encoding target block) of the original image, and outputs it to the orthogonal transform unit 14 as a prediction residual signal.

[0041] The orthogonal transformation unit 14 applies a predetermined orthogonal transformation process to the predicted residual signal input from the predicted residual signal generation unit 13 and generates a signal of the transformation coefficients. For example, this predetermined orthogonal transformation process can be an orthogonal transformation process used in coding schemes such as DCT or DST, or an integer orthogonal transformation process defined in H.264 or H.265, which approximates it to integers, as long as it conforms to the coding scheme being used.

[0042] The image encoding device 1 of the first embodiment performs quantization and entropy encoding on the signal of the conversion coefficients and outputs it externally. In the entropy encoding process, the image can be transmitted by converting it into a code using arithmetic encoding such as CABAC, along with various encoding parameters. The encoding parameters can include selectable parameters such as inter-prediction parameters, intra-prediction parameters, block size parameters related to block division (block division parameters), and quantization parameters related to quantization processing.

[0043] Next, as a typical example of the adjacent pixel non-reference prediction unit 11, an inter-prediction unit 11a using inter-prediction will be described, and the configuration of the image encoding device 1 and an example of its operation, in which the prediction signal of the predicted image block is filtered by the filter processing unit 12, will be described with reference to Figures 3 to 5.

[0044] Figure 3 is a block diagram showing an embodiment of the image coding device 1 according to the first embodiment of the present invention. The image coding device 1 shown in Figure 3 comprises a preprocessing unit 10, an interpretation unit 11a, a filter processing unit 12, a prediction residual signal generation unit 13, an orthogonal transformation unit 14, a quantization unit 15, an inverse quantization unit 16, an inverse orthogonal transformation unit 17, a decoded image generation unit 18, an in-loop filter unit 19, a frame memory 20, an intraprediction unit 21, a motion vector calculation unit 22, an entropy coding unit 23, and a prediction image selection unit 24.

[0045] The preprocessing unit 10 divides the original image of each frame of the input video data into encoding target blocks of a predetermined block size and outputs them to the prediction residual signal generation unit 13 in a predetermined order.

[0046] The prediction residual signal generation unit 13 calculates the error of each pixel signal in the block of the predicted image obtained from the filter processing unit 12 for each pixel signal in the block to be encoded, and outputs it to the orthogonal transform unit 14 as a block-level prediction residual signal.

[0047] The orthogonal transformation unit 14 applies a predetermined orthogonal transformation process to the predicted residual signal input from the predicted residual signal generation unit 13 and generates a signal of the transformation coefficients.

[0048] The quantization unit 15 applies a predetermined quantization process to the conversion coefficient signal obtained from the orthogonal transformation unit 14 and outputs it to the entropy coding unit 23 and the inverse quantization unit 16.

[0049] The inverse quantization unit 16 performs inverse quantization on the quantized conversion coefficient signal obtained from the quantization unit 15 and outputs it to the inverse orthogonal transformation unit 17.

[0050] The inverse orthogonal transform unit 17 reconstructs the predicted residual signal by applying an inverse orthogonal transform process to the inverse quantization conversion coefficient signal obtained from the inverse quantization unit 16, and outputs it to the decoded image generation unit 18.

[0051] The decoded image generation unit 18 adds the predicted residual signal, restored by the inverse orthogonal transform unit 17, to the predicted image block, which is predicted by the intra prediction unit 21 or the inter prediction unit 11a and obtained from the filter processing unit 12, to generate a local decoded image block, which is then output to the in-loop filter unit 19.

[0052] The in-loop filtering unit 19 applies in-loop filtering, such as an Adaptive Loop Filter (ALF), Sample Adaptive Offset (SAO), or deblocking filter, to blocks of the locally decoded image obtained from the decoded image generation unit 18, and outputs them to the frame memory 20. The filter parameters related to these filtering processes are output to the entropy encoding unit 23 as one of the encoding parameters used as supplementary information for the encoding process.

[0053] The frame memory 20 stores blocks of locally decoded images obtained through the in-loop filter unit 19 and holds them as reference images usable by the intra-prediction unit 21, the inter-prediction unit 11a, and the motion vector calculation unit 22.

[0054] The intra prediction unit 21, when the prediction image selection unit 24 selects a signal to predict using only the signals within the frame to be predicted, performs DC prediction, Planar prediction, or directional prediction on the pixel signals (original signals) within the encoding target block of the original image using the pixel signals of encoded and decoded blocks adjacent to the left or above the encoding target block, which are stored as reference images in the frame memory 20. It generates a block of predicted images consisting of the pixel signals (predicted signals) predicted by extrapolation and outputs it to the filter processing unit 12.

[0055] Furthermore, the intra-prediction parameters used by the intra-prediction unit 21 to identify DC prediction, Planer prediction, or direction prediction are output to the entropy coding unit 23 as one of the coding parameters used as supplementary information for the coding process.

[0056] When the prediction image selection unit 24 selects a signal to perform motion compensation prediction between frames for P / B frames, the interpretation unit 11a generates blocks of predicted images by compensating the pixel signals (original signals) in the encoding target blocks of the original image with motion vectors provided by the motion vector calculation unit 22 using data stored as a reference image in the frame memory 20, and outputs these blocks to the filter processing unit 12.

[0057] The motion vector calculation unit 22 searches for the position that is most similar to the encoding target block of the original image using block matching techniques or the like with the reference image data stored in the frame memory 20, calculates a value indicating the spatial shift as a motion vector, and outputs it to the interpretation unit 11a.

[0058] Furthermore, the interprediction parameters used in interprediction, including the motion vector calculated by the motion vector calculation unit 22, are output to the entropy coding unit 23 as one of the coding parameters to be used as supplementary information for the coding process.

[0059] The entropy coding unit 23 applies entropy coding processing to the output signal from the quantization unit 15 and various coding parameters, and outputs a stream signal of the coded video data.

[0060] An example of the operation of the filter processing unit 12 will be explained with reference to Figures 4 and 5. As shown in Figure 4, when the filter processing unit 12 receives a predicted image from the intra prediction unit 21 or the inter prediction unit 11a (step S1), it identifies whether the predicted image is an adjacent pixel non-reference prediction, that is, in this example, whether it is an inter prediction based on signal selection by the predicted image selection unit 24 (step S2). The same applies when the adjacent pixel non-reference prediction is the intra block copy prediction or cross component signal prediction described above.

[0061] Next, if the filter processing unit 12 determines that the predicted image is not an adjacent pixel non-reference prediction, i.e., an intra-prediction in this example (step S2: No), it determines whether or not to perform filtering according to a predetermined processing method (step S5). If filtering is not performed on the predicted image of an intra-prediction (step S5: No), the filter processing unit 12 outputs the predicted residual signal to the prediction residual signal generation unit 13 without performing filtering on the predicted signal (step S6). On the other hand, if filtering is performed on the predicted image of an intra-prediction (step S5: Yes), the filter processing unit 12 proceeds to step S3. Note that filtering on the predicted image of an intra-prediction can be done in the same way as currently defined H.265 and is not directly related to the essence of the present invention, so the explanation of step S3 will describe an example of performing filtering on the predicted image of an inter-prediction.

[0062] If the filter processing unit 12 determines that the predicted image is an adjacent pixel non-reference prediction, that is, in this example, an inter-prediction (Step S2: Yes), it selects the pixel signals (prediction signals) of a predetermined filtering target region from the predicted image (Step S3). For example, as illustrated in Figure 5, the leftmost and uppermost pixel signals (prediction signals) of the predicted image are defined as the filtering target region Sf from among the pixel signals (prediction signals) within a 4x4 block of the predicted image.

[0063] Next, the filter processing unit 12 uses the decoded signals adjacent to the left and above the predicted image (decoded adjacent signals) to apply a low-pass filter to the pixel signals (predicted signals) of the selected filtering target area to generate a new predicted image consisting of new predicted signals, and outputs it to the predicted residual signal generation unit 13 (step S4). For example, as illustrated in Figure 5, the filter processing unit 12 applies a smoothing filter with predetermined weight coefficients to the horizontal filter area, vertical filter area, and angular filter area of ​​the pixel signals (predicted signals) in the 4x4 predicted image block Bkp, respectively. The signal is then processed and output to the predicted residual signal generation unit 13.

[0064] Furthermore, the filtering performed by the filter processing unit 12 can be any low-pass filter, and filtering other than smoothing filters can be applied. The region of the decoded adjacent signal and the weight coefficients used for filtering are not limited to the example shown in Figure 5. For example, in the case of a large-size coded block, filtering may be applied to multiple rows or columns. In addition, the filter type, the region to be filtered, and the weight coefficients related to the filtering performed by the filter processing unit 12 can be configured to be transmitted as supplementary information to the coding process, but they do not need to be transmitted if they are predetermined between the sender and receiver (between coding and decoding).

[0065] By applying a filter to the predicted image using the filter processing unit 12, regardless of the type of orthogonal transformation processing performed by the orthogonal transformation unit 14, the signal error between the leftmost and uppermost pixel signals (predicted signals) in the predicted image and the adjacent decoded signals is reduced, thereby improving the encoding efficiency of the video.

[0066] (Image decoding device) Next, the filter processing unit 54 and its peripheral function blocks, which are the main components of the image decoding device 5 according to the first embodiment of the present invention, will be described with reference to Figure 6, and a specific typical example of the image decoding device 5 will be described with reference to Figure 7.

[0067] Figure 6 is a block diagram of the area around the filter processing unit 54 related to the predicted image in the image decoding device 5 of the first embodiment of the present invention. As shown in Figure 6, the image decoding device 5 of the first embodiment of the present invention is configured to include an inverse orthogonal transform unit 52, an adjacent pixel non-reference prediction unit 53, a filter processing unit 54, and a decoded image generation unit 55.

[0068] The inverse orthogonal transform unit 52 receives the stream signal transmitted from the image encoding device 1, performs an inverse orthogonal transform on the transformation coefficients recovered through entropy decoding and inverse quantization processing, and outputs the resulting predicted residual signal to the decoded image generation unit 55.

[0069] The adjacent pixel non-reference prediction unit 53 is a functional unit corresponding to the adjacent pixel non-reference prediction unit 11 on the image encoding device 1 side, and performs signal prediction processing to generate a predicted image without using adjacent decoded signals (decoded adjacent signals). Adjacent pixel non-reference predictions include, for example, inter-prediction which performs motion compensation prediction between frames, intra-block copy prediction which generates a predicted image by copying different decoded partial images of the same frame, and cross-component signal prediction which generates a predicted image by utilizing the correlation between component signals, namely the luminance signal and the chrominance signal, at corresponding block positions in a given frame.

[0070] The filter processing unit 54 of the first embodiment is a functional unit corresponding to the filter processing unit 12 on the image encoding device 1 side. It uses the decoded signals (decoded adjacent signals) that are adjacent to the left and above the predicted image, respectively, from the pixel signals (predicted signals) within the block of the predicted image generated by the adjacent pixel non-reference prediction unit 53, and applies a low-pass filter to the predicted signals to generate a new predicted image consisting of new predicted signals, which is then output to the decoded image generation unit 55.

[0071] The decoded image generation unit 55 is a functional unit that generates a decoded image block consisting of a decoded signal by adding the predicted residual signal restored by the inverse orthogonal transform unit 52 to the predicted image block obtained from the filter processing unit 54, which will be described later.

[0072] That is, the filter processing unit 12 in the image encoding device 1 filters the predicted image. The predicted residual signal generated after processing is transmitted with improved video encoding efficiency, allowing the image decoding device 5 to efficiently restore the transmitted video. Similarly, in the filter processing unit 54, as with the filter processing unit 12, any low-pass filter is acceptable, and filters other than smoothing filters can be applied. The region and weight coefficients of the decoded adjacent signals used for filtering can also be determined as appropriate. Furthermore, if the filter type, filtering region, and weight coefficients related to filtering by the filter processing unit 12 are predetermined between transmission and reception (encoding and decoding), or transmitted as supplementary information to the encoding process, the filter processing unit 54 of the image decoding device 5 can perform filtering accordingly, thereby improving the accuracy of the decoded signal.

[0073] Next, as a typical example of the adjacent pixel non-reference prediction unit 53, an inter-prediction unit 53a using inter-prediction will be described, and the configuration of the image decoding device 5 and an example of its operation, in which the predicted signal from the inter-prediction is filtered by the filter processing unit 54, will be described with reference to Figure 7.

[0074] Figure 7 is a block diagram showing one embodiment of the image decoding device 5 according to the first embodiment of the present invention. The image decoding device 5 shown in Figure 7 comprises an entropy decoding unit 50, an inverse quantization unit 51, an inverse orthogonal transformation unit 52, an interpretation unit 53a, a filter processing unit 54, a decoded image generation unit 55, an in-loop filter unit 56, a frame memory 57, an intraprediction unit 58, and a predicted image selection unit 59.

[0075] The entropy decoding unit 50 receives the stream signal transmitted from the image encoding device 1 and outputs various encoding parameters obtained by performing entropy decoding processing corresponding to the entropy encoding processing of the image encoding device 1 to each functional block, and also outputs the quantized conversion coefficients to the inverse quantization unit 51. As various encoding parameters, for example, filter parameters, intra-prediction parameters and inter-prediction parameters are output to the in-loop filter unit 56, intra-prediction unit 58, and inter-prediction unit 53a, respectively.

[0076] The inverse quantization unit 51 applies an inverse quantization process corresponding to the quantization process of the image coding device 1 to the quantized conversion coefficients obtained from the entropy decoding unit 50, and outputs the resulting conversion coefficients to the inverse orthogonal transformation unit 52.

[0077] The inverse orthogonal transform unit 52 reconstructs the predicted residual signal by applying an inverse orthogonal transform process corresponding to the orthogonal transform process of the image encoding device 1 to the transformation coefficients obtained from the inverse quantization unit 51, and outputs it to the decoded image generation unit 55.

[0078] The decoded image generation unit 55 generates a decoded image block by adding the predicted residual signal restored by the inverse orthogonal transform unit 52 to the predicted image block generated by the predicted image block, which is predicted by the intra prediction unit 58 or the inter prediction unit 53a and subjected to a low-pass filter by the filter processing unit 54, and outputs it to the in-loop filter unit 56.

[0079] The in-loop filter unit 56 performs a filter process corresponding to the in-loop filter process on the image encoding device 1 side and outputs it to the frame memory 57.

[0080] The frame memory 57 stores blocks of decoded images obtained via the in-loop filter unit 56 and holds them as reference images usable by the intra-prediction unit 58 and the inter-prediction unit 53a.

[0081] When the prediction image selection unit 59 performs signal prediction using only the signals within the frame to be predicted, the intra prediction unit 58 uses the intra prediction parameters and the pixel signals of the decoded blocks adjacent to the left or above the decoded block stored as a reference image in the frame memory 57 to perform the corresponding prediction processing with the intra prediction unit 21 on the image encoding device 1 side to generate a block of predicted images and output it to the filter processing unit 54.

[0082] When the prediction image selection unit 59 selects a signal to perform motion compensation prediction between frames for P / B frames, the interpretation unit 53a uses the motion vector included in the interpretation parameters to compensate for the motion of the data stored as a reference image in the frame memory 20, thereby generating a block of predicted images and outputting it to the filter processing unit 54.

[0083] Furthermore, the system can be configured to construct frames from the decoded data stored in the frame memory 57 and output them as decoded images.

[0084] The filter processing unit 54 operates in the same manner as illustrated with reference to Figures 4 and 5. Specifically, when the predicted image is an adjacent pixel non-reference prediction, i.e., an inter-prediction in this example, the filter processing in the filter processing unit 54 selects a pixel signal (prediction signal) from a predetermined filtering target area of ​​the predicted image, and uses the decoded adjacent signal (decoded adjacent signal) adjacent to the left or above the predicted image to apply a low-pass filter to the selected pixel signal (prediction signal) of the filtering target area to generate a new predicted image consisting of a new prediction signal, which is then output to the decoded image generation unit 55. For example, as illustrated in Figure 5, the filter processing unit 54 uses the decoded adjacent signal So around the block boundary of block Bkp of the predicted image to apply a smoothing filter with predetermined weight coefficients to the horizontal filter area, vertical filter area, and angular filter area of ​​the filtering target area Sf within the pixel signal (prediction signal) of block Bkp of the predicted image, and outputs it to the decoded image generation unit 55.

[0085] According to the image encoding device 1 and image decoding device 5 of the first embodiment configured as described above, the signal error occurring between the predicted signal of the predicted image and the adjacent decoded block signal is reduced, improving encoding efficiency. Therefore, an image encoding device and image decoding device for a video encoding method with high encoding efficiency can be realized. In other words, the residual component in the predicted residual signal can be reduced, and encoding efficiency can be improved by reducing the amount of information to be encoded and transmitted.

[0086] [Second Embodiment] Next, the respective filter processing units 12 and 54 in the image encoding device 1 and image decoding device 5 of the second embodiment of the present invention will be described with reference to Figure 8. In the filter processing units 12 and 54 of the first embodiment, an example was described in which a new predicted image is generated by applying a low-pass filter to the predicted signal using the decoded signals (decoded adjacent signals) that are adjacent to the left and above the predicted image, respectively, from among the pixel signals (predicted signals) in the block of the predicted image generated by adjacent pixel non-reference prediction.

[0087] On the other hand, the filter processing units 12 and 54 of this embodiment are configured to perform a correlation determination to determine whether or not to apply a low-pass filter to the predicted signal targeted for the low-pass filter, using the decoded adjacent signal referenced in the low-pass filter processing, among the pixel signals (predicted signals) within the block of the predicted image generated by adjacent pixel non-reference prediction. The determination is made whether or not to apply the low-pass filter processing according to the result of this correlation determination. Similar components will be described using the same reference number. Therefore, the filter processing units 12 and 54 of this embodiment can be applied to the image encoding device 1 and image decoding device 5 illustrated in Figures 3 and 7, respectively. Here, only the processing content related to the filter processing units 12 and 54 of this embodiment will be described, and further detailed explanations will be omitted.

[0088] Figure 8 is a block diagram of the filter processing units 12 and 54 in the image encoding device 1 or image decoding device 5 of this embodiment. Figure 9 is a flowchart relating to the filter processing 12 and 54 of the predicted image in the image encoding device 1 or image decoding device 5 of this embodiment.

[0089] The filter processing units 12 and 54 of this embodiment include a horizontal correlation determination unit 101, a vertical correlation determination unit 102, a filter processing determination unit 103, and a filter processing execution unit 104. The operation of these functional blocks will be explained with reference to the processing example shown in Figure 9. Although the filter processing units 12 and 54 of this embodiment can be configured to process without distinguishing whether it is an adjacent pixel non-reference prediction or not, for example, whether it is an inter-prediction or an intra-prediction, here we will explain an example in which filtering is applied to a predicted image of adjacent pixel non-reference prediction according to the present invention.

[0090] When the filter processing units 12 and 54 receive a predicted image (step S11), the horizontal correlation determination unit 101 and the vertical correlation determination unit 102 each perform a correlation determination process for each filtered area that is individually defined (step S12).

[0091] Specifically, the horizontal correlation determination unit 101 determines the correlation between the predicted image and the decoded adjacent signal adjacent to the left (step S13). If it determines that the horizontal correlation is high (step S13: Yes), it decides to perform horizontal filtering using the decoded adjacent signal (step S14). If it determines that the horizontal correlation is low (step S13: No), it outputs to the filtering determination unit 103 that horizontal filtering will not be performed.

[0092] Similarly, the vertical correlation determination unit 102 determines the correlation between the predicted image and the decoded adjacent signal adjacent to the upper side (step S15). If it determines that the vertical correlation is high (step S15: Yes), it decides to perform vertical filtering using the decoded adjacent signal (step S16). If it determines that the vertical correlation is low (step S15: No), it outputs to the filtering determination unit 103 that vertical filtering will not be performed.

[0093] The filter processing determination unit 103 determines whether it has received a decision to perform both horizontal and vertical filtering (step S17). If it has received a decision to perform both horizontal and vertical filtering (step S17: Yes), it decides to perform both horizontal and vertical filtering and a 3-tap corner filtering (step S18). If it has received a decision not to perform either or both of the horizontal and vertical filtering (step S17: No), it decides to perform filtering only in the horizontal direction when the horizontal correlation is high, and filtering only in the vertical direction when the vertical correlation is high. In addition, if only one of the horizontal or vertical filtering is performed, it decides to perform either horizontal or vertical corner filtering (step S19).

[0094] The processing order from steps S13 to S19 shown in Figure 9 is merely an example and can be configured in any order. Therefore, the filter processing determination unit 103 performs a correlation determination for each of the individually defined filtering target areas in the horizontal and vertical directions, and decides whether to perform the filtering "horizontally only," "vertically only," "horizontally, vertically, and angularly," or "not to perform filtering in any of the horizontal, vertically, or angularly," and outputs this decision to the filter processing execution unit 104.

[0095] The filter processing execution unit 104 performs filtering according to the judgment result of the filter processing determination unit 103 and outputs the predicted image after filtering (step S20). If the judgment result of the filter processing determination unit 103 is "Do not perform on horizontal, vertical, or angle filters", it outputs the predicted image input to the filter processing units 12 and 54 as is.

[0096] For example, the horizontal correlation determination unit 101 calculates the difference between adjacent pixels for each of the four horizontal determination regions A1 to A4 shown in Figure 10(a). If the sum of these differences in the horizontal determination regions A1 to A4 is less than or equal to a predetermined threshold, it determines that the correlation is high; otherwise, it determines that the correlation is low. Similarly, the vertical correlation determination unit 102 calculates the difference between adjacent pixels for each of the four vertical determination regions B1 to B4 shown in Figure 10(b). If the sum of these differences in the vertical determination regions B1 to B4 is less than or equal to a predetermined threshold, it determines that the correlation is high; otherwise, it determines that the correlation is low. By applying a filter in the direction of high correlation as shown in Figure 5, the filter can be set to "execute only in the horizontal direction," "execute only in the vertical direction," "execute in all directions (horizontal, vertical, and angle)," or "not execute in any direction (horizontal, vertical, or angle)." When performing a filter in only the horizontal or vertical direction, the horizontal or vertical filter is applied to the pixels (prediction signals) located at the corners of the block of the prediction image.

[0097] While it is not necessary for the correlation determination area and the filtering area to be the same, doing so simplifies the processing. Furthermore, it should be noted that various techniques are possible for the specific methods of correlation determination and filtering, and this is merely one example.

[0098] By configuring the filter processing units 12 and 54 as in this embodiment and applying them to the image encoding device 1 and the image decoding device 5, respectively, it is possible to improve encoding efficiency while suppressing the resulting degradation of image quality.

[0099] [Third Embodiment] (Image encoding device) Next, the filter processing unit 12 and the orthogonal transform selection control unit 25, which are the main components of the present invention in the image encoding device 1 of the third embodiment of the present invention, will be described with reference to Figures 1(b) and 11, and two specific embodiments of the image encoding device 1 will be described with reference to Figures 5, 12 to 15. The following description will focus on the differences from the embodiments described above. That is, in this embodiment, the same reference numerals are used for components that are the same as in the embodiments described above, and further detailed descriptions thereof will be omitted.

[0100] Figure 11(a) is a block diagram of the filter processing unit 12 and the orthogonal transform selection control unit 25 related to the predicted image in the image encoding device 1 of this embodiment, and Figure 1(b) is an explanatory diagram showing an example of filtering the predicted image by the filter processing unit 12.

[0101] As shown in Figure 11(a), the image encoding device 1 of this embodiment is configured to include an adjacent pixel non-reference prediction unit 11, a filter processing unit 12, a prediction residual signal generation unit 13, an orthogonal transformation unit 14, and an orthogonal transformation selection control unit 25.

[0102] In this embodiment, the filter processing unit 12 is a functional unit that, under the control of the orthogonal transformation selection control unit 25, uses the decoded signals (decoded adjacent signals) that are adjacent to the left and above the predicted image, respectively, from among the pixel signals (predicted signals) in the block of the predicted image generated by the adjacent pixel non-reference prediction unit 11, and applies a low-pass filter to the predicted signal to generate a new predicted image consisting of a new predicted signal, which is then output to the predicted residual signal generation unit 13.

[0103] The orthogonal transformation unit 14, under the control of the orthogonal transformation selection control unit 25, applies one of two types of orthogonal transformation processing to the predicted residual signal input from the predicted residual signal generation unit 13, and generates a signal of the transformation coefficient.

[0104] For example, these two types of orthogonal transformations can be either orthogonal transformations with closed ends in the transformation basis (e.g., DST (Discrete Sine Transform)) or orthogonal transformations with open ends in the transformation basis (e.g., DCT (Discrete Cosine Transform)). There is no limitation on whether the precision should be real or integer, but it can be an integer orthogonal transformation as defined in H.264 or H.265, which approximate integers, as long as it conforms to the encoding scheme used.

[0105] Here, examples of basis waveforms for orthogonal transformations are shown in Figures 11(b) and 11(c). The examples shown in Figures 11(b) and 11(c) illustrate the patterns of the cosine and sine frequency components in the DCT transformation basis (Figure 11(b)) and DST transformation basis (Figure 11(c)), respectively, for four transformation basis points (N=4) that can be used when orthogonal transforming a 4x4 block of predicted residual signals, and illustrate the transformation basis waveforms from low to high frequencies (u=0 to 3 for DCT, and u=1 to 4 for DST). As shown in Figures 11(b) and 11(c), the main difference other than u=0 for DCT and u=4 for DST is the phase, and it can be seen that each transformation basis for a corresponding frequency (e.g., u=2) has the same inter-pixel correlation (same basis amplitude) at the same frequency, but its phase is shifted by π / 2. In both transformed basis waveforms, DCT has large values ​​at its endpoints, resulting in an open circuit, as shown in Figure 11(b), thus achieving an orthogonal transform with open endpoints in the transformed basis. On the other hand, DST has small values ​​at its endpoints, resulting in a closed circuit, as shown in Figure 11(c), thus achieving an orthogonal transform with closed endpoints in the transformed basis.

[0106] Furthermore, "orthogonal transformation processing with closed ends of the transformation basis" refers to a case where one end of the predicted residual signal block that is adjacent to the predicted reference block is closed. For example, as shown in Figure 11(c), an asymmetric DST type 7 transformation basis with the other end open may also be used.

[0107] When predicting adjacent pixels without reference, the orthogonal transformation selection control unit 25, based on predetermined evaluation criteria, instructs the orthogonal transformation unit 14 to use an orthogonal transformation process with closed ends in the transformation basis (e.g., DST) for the blocks of predicted residual signals obtained from the prediction residual signal generation unit 13 based on the blocks of the predicted image after the low-pass filtering. When the low-pass filtering is not performed on the blocks of the predicted image based on the predetermined evaluation criteria, the orthogonal transformation selection control unit 25 instructs the orthogonal transformation unit 14 to use an orthogonal transformation process with open ends in the transformation basis (e.g., DCT) for the blocks of predicted residual signals.

[0108] Two embodiments will be described as predetermined evaluation criteria for the orthogonal transformation selection control unit 25.

[0109] As will be described in detail later, in the first embodiment, the orthogonal transformation selection control unit 25 instructs the filter processing unit 12 to select a signal for the filtering target region as shown in Figure 1(b), and to perform a preliminary execution of low-pass filtering using the decoded adjacent signals. The filter processing unit 12 then obtains each predicted image with and without filtering as predicted block information (est1). Subsequently, the orthogonal transformation selection control unit 25 compares each predicted image with and without filtering from this predicted block information (est1). If it is determined that the influence of the endpoints of the prediction region (i.e., the predicted residual signals for the pixel signals located on the outer periphery of the block of the predicted image) is relatively large relative to the block size (in this example, 4x4 is used as an example, but the block size is not limited), the filter processing unit 12 outputs the predicted image after the filtering is performed to the predicted residual signal generation unit 13. Then, the orthogonal transformation selection control unit 25 controls the orthogonal transformation unit 14 to use an orthogonal transformation process with closed ends for the transformation basis (for example, DST) for the block of predicted residual signal obtained from the predicted residual signal generation unit 13, based on the block of predicted image after low-pass filtering.

[0110] On the other hand, the orthogonal transformation selection control unit 25 compares the predicted images obtained from the filter processing unit 12, both those with and without the filter processing, and determines that the influence of the endpoints of the prediction region (i.e., the predicted residual signals for the pixel signals located on the outer periphery of the block of the predicted image) is relatively small with respect to the block size (in this example, 4x4 is used, but the block size is not limited), and causes the filter processing unit 12 to output the predicted image without the filter processing to the predicted residual signal generation unit 13. The orthogonal transformation selection control unit 25 then controls the orthogonal transformation unit 14 to use an orthogonal transformation process with open ends in the transformation basis (for example, DCT) for the predicted residual signal block obtained from the predicted residual signal generation unit 13 based on the predicted image block without the low-pass filter processing.

[0111] Here, various evaluation methods can be envisioned to determine whether the influence of the endpoints of the prediction region is relatively large relative to the block size (in this example, 4x4 is used, but the block size is not limited). For example, it can be determined by comparing the ratio of the variance of the density distribution of the endpoints of the prediction region to the variance of the density distribution of the entire block size and determining whether it is above a predetermined level. Thus, in the first embodiment, the orthogonal transformation selection control unit 25 determines whether or not to perform the filtering process based on the determination criterion of whether or not the filtering effect of the filtering process is above a predetermined level, and further determines the type of orthogonal transformation process to apply according to whether or not the filtering process is performed, and controls the filtering processing unit 12 and the orthogonal transformation unit 14.

[0112] In the second embodiment, the orthogonal transform selection control unit 25 instructs the filter processing unit 12 to select a signal in the region to be filtered as shown in Figure 1(b), and to perform a preliminary low-pass filter processing using the decoded adjacent signal. The unit then acquires and compares rate distortion information (est2) of the generated coding amount (R) and coding distortion amount (D) in two cases: when the orthogonal transform is performed using an orthogonal transform with closed ends of the transform basis (e.g., DST) and then quantization and entropy coding are performed, and when the orthogonal transform is performed using an orthogonal transform with open ends of the transform basis (e.g., DCT) and then quantization and entropy coding are performed, and the unit selects the combination that is superior in terms of RD cost (for example, the combination of performing the filter processing and using DST, and the combination of not performing the filter processing and using DCT).

[0113] Through this RD optimization, if the combination of "execution of filtering and orthogonal transformation with closed ends of the transformation basis" is superior, the orthogonal transformation selection control unit 25 outputs the predicted image after the filtering process from the filtering processing unit 12 to the predicted residual signal generation unit 13. Based on this block of predicted image after low-pass filtering, the control unit 25 instructs the orthogonal transformation unit 14 to use an orthogonal transformation with closed ends of the transformation basis (for example, DST) for the block of predicted residual signals obtained from the predicted residual signal generation unit 13.

[0114] On the other hand, if the combination of "non-performing filtering and orthogonal transformation processing with open ends in the transformation basis" is superior, the orthogonal transformation selection control unit 25 causes the filtering processing unit 12 to output a predicted image for which filtering has not been performed to the predicted residual signal generation unit 13, and controls the orthogonal transformation unit 14 to use an orthogonal transformation processing with open ends in the transformation basis (for example, DCT) for the block of predicted residual signals obtained from the predicted residual signal generation unit 13 based on the block of predicted image for which low-pass filtering has not been performed.

[0115] Thus, in the second embodiment, the orthogonal transformation selection control unit 25 determines whether or not to perform filtering based on RD optimization, and further determines the type of orthogonal transformation processing to apply depending on whether or not filtering is performed, and controls the filtering processing unit 12 and the orthogonal transformation unit 14.

[0116] Finally, the image encoding device 1 of this embodiment performs quantization and entropy encoding on the conversion coefficient signal obtained from the orthogonal transformation unit 14 and outputs it externally. In the entropy encoding process, the signal is converted into a code using arithmetic encoding such as CABAC, along with various encoding parameters, and the image can be transmitted. In addition to the conversion type identification signal according to the present invention, the encoding parameters can also include selectable parameters such as inter-prediction parameters, intra-prediction parameters, block size parameters related to block division (block division parameters), and quantization parameters related to quantization processing.

[0117] Below, we will describe, in more detail, the configuration examples of the image encoding device 1 in which the interpretation unit 11a using interpretation is used as a typical example of the adjacent pixel non-reference prediction unit 11, and the above-described embodiments are applied as predetermined evaluation criteria in the orthogonal transformation selection control unit 25.

[0118] (First embodiment) Figure 12 is a block diagram showing a first embodiment of the image coding device 1 of this embodiment. The image coding device 1 shown in Figure 12 comprises a preprocessing unit 10, an interpretation unit 11a, a filter processing unit 12, a prediction residual signal generation unit 13, an orthogonal transformation unit 14, a quantization unit 15, an inverse quantization unit 16, an inverse orthogonal transformation unit 17, a decoded image generation unit 18, an in-loop filter unit 19, a frame memory 20, an intraprediction unit 21, a motion vector calculation unit 22, an entropy coding unit 23, a prediction image selection unit 24, and an orthogonal transformation selection control unit 25.

[0119] The orthogonal transformation unit 14, under the control of the orthogonal transformation selection control unit 25, applies a predetermined orthogonal transformation process to the predicted residual signal input from the predicted residual signal generation unit 13 and generates a signal of the transformation coefficients.

[0120] In the first embodiment, the filter processing unit 12 controls the execution or non-execution of low-pass filtering under the control of the orthogonal transformation selection control unit 25. When low-pass filtering is performed, the filter processing unit 12 uses the decoded signals (decoded adjacent signals) that are adjacent to the left and above the predicted image, respectively, from the pixel signals (predicted signals) within the block of the predicted image generated by the inter-prediction unit 11a, which is an adjacent-pixel non-reference prediction, to apply low-pass filtering to the predicted signal, thereby generating a new predicted image consisting of a new predicted signal, which is output to the prediction residual signal generation unit 13. When low-pass filtering is not performed, the predicted image generated by the inter-prediction unit 11a is output to the prediction residual signal generation unit 13. In the case of predictions other than adjacent-pixel non-reference predictions, such as intra-prediction, the filter processing unit 12 shall follow a predetermined method for whether or not to execute filtering.

[0121] The operation of the orthogonal transformation selection control unit 25 in the first embodiment will be described with reference to Figure 13. Figure 13 is a flowchart of the processing of the orthogonal transformation selection control unit 25 in the first embodiment of the image encoding device 1 of this embodiment. First, the orthogonal transformation selection control unit 25 receives notification from, for example, the filter processing unit 12 or the prediction image selection unit 24, and determines whether the prediction image to be processed by the filter processing unit 12 is an adjacent pixel non-reference prediction (step S1).

[0122] When predicting adjacent pixels without reference (Step S1: Yes), the orthogonal transformation selection control unit 25 instructs the filter processing unit 12 to select a signal for the region to be filtered and to perform a preliminary low-pass filter processing using the decoded adjacent signal (Step S2).

[0123] For example, as illustrated in Figure 5, the filter processing unit 12 defines the leftmost and uppermost pixel signals (prediction signals) of the 4x4 predicted image block as the filtering target region Sf. Subsequently, the filter processing unit 12 uses the decoded signals adjacent to the left and above the predicted image (decoded adjacent signals) to apply a low-pass filter to the selected filtering target region's pixel signals (prediction signals) to generate a new predicted image consisting of new prediction signals. For example, as illustrated in Figure 5, the filter processing unit 12 applies a smoothing filter with predetermined weight coefficients to the horizontal filter region, vertical filter region, and angular filter region of the filtering target region Sf, among the pixel signals (prediction signals) in the 4x4 predicted image block Bkp, to generate the filtered predicted image.

[0124] Furthermore, the filtering performed by the filter processing unit 12 can be any low-pass filter, and filtering other than smoothing filters can be applied. The region of the decoded adjacent signal and the weight coefficients used for filtering are not limited to the example shown in Figure 5. For example, in the case of a large-size coded block, filtering may be applied to multiple rows or columns. In addition, the filter type, the region to be filtered, and the weight coefficients related to the filtering performed by the filter processing unit 12 can be configured to be transmitted as supplementary information to the coding process, but they do not need to be transmitted if they are predetermined between the sender and receiver (between coding and decoding).

[0125] Next, referring to Figure 13, the orthogonal transformation selection control unit 25 acquires each predicted image with and without the filter processing from the filter processing unit 12 as predicted block information (est1) (step S3).

[0126] Next, the orthogonal transformation selection control unit 25 compares the predicted images with and without the filtering process based on the predicted block information (est1), and determines whether the effect of the filtering process is above a predetermined level based on whether the influence of the endpoints of the predicted region (i.e., the predicted residual signals for the pixel signals located on the outer periphery of the block of the predicted image) is relatively large relative to the block size (in this example, 4x4 is used, but the block size is not limited) (step S4).

[0127] When it is determined that the influence of the endpoints of the prediction region is relatively large relative to the block size (Step S4: Yes), the orthogonal transformation selection control unit 25 causes the filter processing unit 12 to output the predicted image after the filtering process to the prediction residual signal generation unit 13, and instructs the orthogonal transformation unit 14 to use an orthogonal transformation process with closed ends for the transformation basis (for example, DST) for the predicted residual signal obtained from the prediction residual signal generation unit 13 based on the block of the predicted image after the low-pass filtering process (Step S5).

[0128] On the other hand, when it is determined that the influence of the endpoints of the prediction region is relatively small relative to the block size (Step S4: No), the orthogonal transformation selection control unit 25 causes the filter processing unit 12 to output the prediction residual signal generation unit 13 with the prediction images for which the filtering process has not been performed, and instructs the orthogonal transformation unit 14 to use an orthogonal transformation process with open ends of the transformation basis (for example, DCT) for the prediction residual signal block obtained from the prediction residual signal generation unit 13 based on the block of prediction images for which the low-pass filtering process has not been performed (Step S6).

[0129] Furthermore, when performing predictions other than adjacent pixel non-reference predictions, such as intra predictions (Step S1: No), the orthogonal transformation selection control unit 25 controls the filter processing unit 12 and the orthogonal transformation unit 14 to perform filtering and the type of orthogonal transformation according to a predetermined method (Step S7).

[0130] Thus, in the first embodiment, the orthogonal transform selection control unit 25 controls the filter processing unit 12 and the orthogonal transform unit 14 to determine whether or not to perform the filter processing and the type of orthogonal transform processing to apply, based on a determination criterion of whether or not the filtering effect of the filtering processing is above a predetermined level. The orthogonal transform selection control unit 25 then outputs a transformation type identification signal indicating which of the two types of orthogonal transform processing was applied to the decoding side via the entropy encoding unit 23 as one of the encoding parameters.

[0131] As described above, the filter processing unit 12 and the orthogonal transformation unit 14 are controlled based on a criterion for determining whether the filtering effect of the filtering process is above a predetermined level, thereby determining whether or not to perform filtering and the type of orthogonal transformation process to apply. This makes it possible to reduce the residual component in the predicted residual signal and improve coding efficiency by reducing the amount of information to be encoded and transmitted.

[0132] In particular, even when predicting without referring to adjacent pixels, the prediction residual signal block obtained based on the block of the predicted image after the filtering process utilizes the correlation between adjacent pixel signals and employs an orthogonal transformation process with closed ends for the transformation basis (e.g., DST). This matches not only the phase but also the characteristics of the transformation basis corresponding to the endpoints of the prediction region, thereby significantly improving coding efficiency.

[0133] (Second example) Next, a second embodiment will be described. Figure 14 is a block diagram showing a second embodiment of the image encoding device 1 of this embodiment. In Figure 14, components similar to those in Figure 12 are given the same reference numerals, and further detailed explanations thereof are omitted.

[0134] In the second embodiment, the filter processing unit 12, similar to the first embodiment, controls the execution or non-execution of low-pass filtering under the control of the orthogonal transformation selection control unit 25. When the filtering is executed, it uses the decoded signals (decoded adjacent signals) that are adjacent to the left and above the predicted image, respectively, from the pixel signals (predicted signals) within the block of the predicted image generated by the inter-prediction unit 11a, which is an adjacent-pixel non-reference prediction, to apply low-pass filtering to the predicted signal, thereby generating a new predicted image consisting of the predicted signal and outputting it to the prediction residual signal generation unit 13. When the filtering is not executed, the predicted image generated by the inter-prediction unit 11a is output to the prediction residual signal generation unit 13. In the case of predictions other than adjacent-pixel non-reference predictions, such as intra-prediction, the filter processing unit 12 shall follow a predetermined method for whether or not to execute filtering.

[0135] The operation of the orthogonal transformation selection control unit 25 in the second embodiment will be described with reference to Figure 15. Figure 15 is a flowchart of the processing of the orthogonal transformation selection control unit 25 in the second embodiment of the image encoding device 1 of this embodiment. First, the orthogonal transformation selection control unit 25 receives notification from, for example, the filter processing unit 12 or the prediction image selection unit 24, and determines whether the prediction image to be processed by the filter processing unit 12 is an adjacent pixel non-reference prediction (step S11).

[0136] When predicting adjacent pixels that do not reference each other (step S11: Yes), the orthogonal transformation selection control unit 25 instructs the filter processing unit 12 to select a signal for the filtering target region as shown in Figure 1(b), and to perform a preliminary low-pass filter processing using the decoded adjacent signal (step S12).

[0137] Next, when the filter processing of the filter processing unit 12 is executed, the orthogonal transform selection control unit 25 outputs the predicted image after the execution of the filter processing to the predicted residual signal generation unit 13 using an orthogonal transform processing (for example, DST) with closed ends of the transformation basis to the orthogonal transform unit 14. Based on this block of predicted image after low-pass filtering, the block of predicted residual signal obtained from the predicted residual signal generation unit 13 is orthogonally transformed, and quantization processing and entropy coding processing are performed by the quantization unit 15 and the entropy coding unit 23. Alternatively, when the filter processing of the filter processing unit 12 is not executed, the orthogonal transform selection control unit 25 outputs the predicted image after the execution of the filter processing to the orthogonal transform unit 14. Using an orthogonal transformation process with open ends (e.g., DCT), the filter processing unit 12 outputs the unprocessed predicted image for the filter processing to the predicted residual signal generation unit 13. Based on the unprocessed predicted image block for the low-pass filter processing, the block of predicted residual signal obtained from the predicted residual signal generation unit 13 is orthogonally transformed, and the rate distortion information (est2) of the generated coding amount (R) and coding distortion amount (D) when quantization processing and entropy coding processing are performed by the quantization unit 15 and entropy coding unit 23 is obtained from the quantization unit 15 and entropy coding unit 23, respectively (step S13).

[0138] Next, the orthogonal transformation selection control unit 25 compares and determines which combination (for example, the combination of performing filtering and DST versus the combination of not performing filtering and DCT) is superior in terms of RD cost based on this rate distortion information (est2) (step S14).

[0139] If it is determined that performing the filtering process and using an orthogonal transformation process with closed ends for the transformation basis (e.g., DST) is superior in terms of RD cost (step S14: Yes), the filtering processing unit 12 outputs the predicted image after the filtering process to the predicted residual signal generation unit 13, and the orthogonal transformation unit 14 is instructed to use an orthogonal transformation process with closed ends for the transformation basis (e.g., DST) for the block of predicted residual signals obtained from the predicted residual signal generation unit 13 based on this block of predicted image after low-pass filtering (step S15).

[0140] On the other hand, if it is determined that not performing the filtering process and using an orthogonal transformation process with open ends for the transformation basis (e.g., DCT) is superior in terms of RD cost (step S14: No), the filtering processing unit 12 outputs the predicted image for which the filtering process has not been performed to the predicted residual signal generation unit 13, and the orthogonal transformation unit 14 is instructed to use an orthogonal transformation process with open ends for the transformation basis (e.g., DCT) for the block of predicted image for which the low-pass filtering process has not been performed, based on the block of predicted image (step S16).

[0141] Furthermore, when performing predictions other than adjacent pixel non-reference predictions, such as intra prediction (step S11: No), the orthogonal transformation selection control unit 25 controls the filter processing unit 12 and the orthogonal transformation unit 14 to perform filtering and the type of orthogonal transformation according to a predetermined method (step S17).

[0142] Thus, in the second embodiment, the orthogonal transformation selection control unit 25 determines whether or not to perform filtering and the type of orthogonal transformation processing to apply based on the results selected by RD optimization, and controls the filtering processing unit 12 and the orthogonal transformation unit 14. The orthogonal transformation selection control unit 25 then outputs a transformation type identification signal indicating which of the two types of orthogonal transformation processing was applied as one of the encoding parameters to the decoding side via the entropy encoding unit 23.

[0143] As described above, the orthogonal transformation selection control unit 25 determines whether or not to perform filtering and the type of orthogonal transformation processing to apply based on the results selected by RD optimization, and controls the filtering processing unit 12 and the orthogonal transformation unit 14. This makes it possible to reduce the residual component in the predicted residual signal and improve coding efficiency by reducing the amount of information to be encoded and transmitted.

[0144] In particular, even when predicting without referencing adjacent pixels, the filtering process is allowed to be performed. Based on the blocks of predicted images after the filtering process, the blocks of predicted residual signals obtained utilize the correlation between adjacent pixel signals and employ an orthogonal transform process with closed ends for the transform basis (e.g., DST). This allows for a more significant improvement in coding efficiency by matching not only the phase but also the characteristics of the transform basis corresponding to the endpoints of the prediction region.

[0145] (Image decoding device) Next, the peripheral function blocks of the filter processing unit 54 and the inverse orthogonal transform selection control unit 60, which are the main components of the image decoding device 5 according to the present invention, will be described with reference to Figure 16, and a specific typical example of the image decoding device 5 will be described with reference to Figure 17.

[0146] Figure 16 is a block diagram of the area around the filter processing unit 54 and the inverse orthogonal transform selection control unit 60 related to the predicted image in the image decoding device 5 of this embodiment. As shown in Figure 16, the image decoding device 5 of this embodiment is configured to include an inverse orthogonal transform unit 52, an adjacent pixel non-reference prediction unit 53, a filter processing unit 54, a decoded image generation unit 55, and an inverse orthogonal transform selection control unit 60.

[0147] The inverse orthogonal transform unit 52, under the control of the inverse orthogonal transform selection control unit 60, receives the stream signal transmitted from the image encoding device 1, performs an inverse orthogonal transform on the transformation coefficients restored through entropy decoding and inverse quantization processing, and outputs the resulting predicted residual signal to the decoded image generation unit 55.

[0148] The filter processing unit 54 is a functional unit corresponding to the filter processing unit 12 on the image encoding device 1 side. Under the control of the inverse orthogonal transform selection control unit 60, the execution or non-execution of low-pass filtering is controlled. When it is executed, the filter processing unit 54 uses the decoded signals (decoded adjacent signals) that are adjacent to the left and above the predicted image, respectively, from the pixel signals (predicted signals) within the block of the predicted image generated by the adjacent pixel non-reference prediction unit 53, to apply low-pass filtering to the predicted signal, thereby generating a new predicted image consisting of the new predicted signal, which is output to the decoded image generation unit 55. When it is not executed, the predicted image generated by the adjacent pixel non-reference prediction unit 53 is output to the decoded image generation unit 55. In the case of predictions other than adjacent pixel non-reference prediction, such as intra prediction, the filter processing unit 54 corresponds to the encoding side, and whether or not filtering is executed follows a predetermined method.

[0149] When predicting adjacent pixels not to be referenced, the inverse orthogonal transform selection control unit 60 refers to the transformation type identification signal obtained from the image encoding device 1 and, when it indicates that an orthogonal transform process with closed ends of the transformation basis (e.g., DST) was used for the transformation coefficients to be processed, it instructs the inverse orthogonal transform unit 52 to apply an inverse orthogonal transform process with closed ends of the transformation basis (e.g., IDST) to the restored transformation coefficients, thereby performing selection control, and also instructs the filter processing unit 54 to execute the corresponding low-pass filter process.

[0150] On the other hand, when the image encoding device 1 provides a conversion type identification signal and indicates that an orthogonal transformation process with open ends in the conversion basis (e.g., DCT) has been used, the inverse orthogonal transformation selection control unit 60 selects the conversion by instructing the inverse orthogonal transformation unit 52 to apply the inverse orthogonal transformation process with open ends in the conversion basis (e.g., IDCT) to the restored conversion coefficients, and also instructs the filter processing unit 54 to not execute the corresponding low-pass filter process.

[0151] In other words, the predicted residual signal generated by the filter processing unit 12 in the image encoding device 1 after filtering the predicted image is transmitted in a state where the encoding efficiency of the video is improved by undergoing an orthogonal transformation process (e.g., DST) with closed ends of the transformation basis, and the image decoding device 5 can efficiently restore the video transmitted in this state. In addition, in the filter processing unit 54, as in the case of the filter processing unit 12, the filtering process can be any low-pass filter, and filters other than smoothing filters can be applied, and the region and weight coefficients of the decoded adjacent signals used for filtering can be determined as appropriate. Furthermore, if the filter type, the region to be filtered, and the weight coefficients related to the filtering process by the filter processing unit 12 are predetermined between the sender and receiver (between encoding and decoding) or transmitted as supplementary information to the encoding process, the filter processing unit 54 of the image decoding device 5 can also perform filtering in the same way according to this information, thereby improving the accuracy of the decoded signal.

[0152] Next, as a typical example of the adjacent pixel non-reference prediction unit 53, we will describe the inter-prediction unit 53a using inter-prediction, and then we will describe an example of the configuration of the image decoding device 5 in which an inverse orthogonal transform selection control unit 60 that selects and controls the inverse orthogonal transform unit 52 and the filter processing unit 54 is applied to the prediction signal obtained by the inter-prediction.

[0153] Figure 17 is a block diagram showing one embodiment of the image decoding device 5 of this embodiment. The image decoding device 5 shown in Figure 17 comprises an entropy decoding unit 50, an inverse quantization unit 51, an inverse orthogonal transform unit 52, an inter-prediction unit 53a, a filter processing unit 54, a decoded image generation unit 55, an in-loop filter unit 56, a frame memory 57, an intra-prediction unit 58, a predicted image selection unit 59, and an inverse orthogonal transform selection control unit 60.

[0154] The inverse orthogonal transform unit 52, under the control of the inverse orthogonal transform selection control unit 60, reconstructs the predicted residual signal by applying an inverse orthogonal transform process corresponding to the orthogonal transform process of the image encoding device 1 to the transformation coefficients obtained from the inverse quantization unit 51, and outputs it to the decoded image generation unit 55.

[0155] The filter processing unit 54 operates under the control of the inverse orthogonal transform selection control unit 60, in the same manner as illustrated with reference to Figure 5. Specifically, the filter processing in the filter processing unit 54 according to this embodiment, when the predicted image is an adjacent pixel non-reference prediction, i.e., an inter-prediction in this example, selects a pixel signal (prediction signal) from a predetermined filtering target area from the predicted image, and uses the decoded signals adjacent to the left or above the predicted image (decoded adjacent signals) to apply a low-pass filter to the selected pixel signal (prediction signal) of the filtering target area to generate a new predicted image consisting of a new prediction signal, which is output to the decoded image generation unit 55. For example, as illustrated in Figure 5, the filter processing unit 54, under the control of the inverse orthogonal transform selection control unit 60, uses the decoded adjacent signal So around the block boundary of block Bkp of the predicted image to perform smoothing filtering on the horizontal filter region, vertical filter region, and angular filter region of the filter target region Sf within the pixel signal (predicted signal) of block Bkp of the predicted image, using predetermined weight coefficients, and outputs the result to the decoded image generation unit 55.

[0156] The operation of the inverse orthogonal transformation selection control unit 60 will be explained with reference to Figure 18. Figure 18 is a flowchart of the processing of the inverse orthogonal transformation selection control unit 60 in one embodiment of the image decoding device 5 of this embodiment. First, the inverse orthogonal transformation selection control unit 60 receives notification from the prediction image selection unit 59 to determine whether the current processing state is adjacent pixel non-reference prediction (step S21).

[0157] Next, when predicting the non-reference of adjacent pixels (step S21: Yes), the inverse orthogonal transform selection control unit 60 refers to the transform type identification signal obtained from the image coding device 1 and determines whether an orthogonal transform process with closed ends of the transform basis (e.g., DST) or an orthogonal transform process with open ends of the transform basis (e.g., DCT) was used with respect to the transform coefficients (step S22).

[0158] Next, when the inverse orthogonal transform selection control unit 60 determines that an orthogonal transform process with closed ends for the transformation basis (e.g., DST) has been used for the transformation coefficients (step S23: Yes), it instructs the filter processing unit 54 to execute the low-pass filter process and also instructs the inverse orthogonal transform unit 52 to apply the inverse orthogonal transform process with closed ends for the transformation basis (e.g., IDST) to the transformation coefficients restored via the inverse quantization unit 51, thereby performing selection control and causing the inverse orthogonal transform unit 52 to generate a predicted residual signal to be added to the predicted image obtained via the filter processing unit 54 (step S24).

[0159] On the other hand, when it is determined that an orthogonal transformation process (e.g., DCT) with open ends of the transformation basis has been used for the transformation coefficients (step S23: No), the inverse orthogonal transformation selection control unit 60 instructs the filter processing unit 54 not to perform the low-pass filter processing and instructs the inverse orthogonal transformation unit 52 to apply the inverse orthogonal transformation process (e.g., IDCT) with open ends of the transformation basis to the transformation coefficients restored via the inverse quantization unit 51, thereby performing selection control and causing the inverse orthogonal transformation unit 52 to generate a predicted residual signal to be added to the predicted image obtained via the filter processing unit 54 (step S25).

[0160] Furthermore, when performing predictions other than adjacent pixel non-reference predictions, such as intra prediction (step S21: No), the inverse orthogonal transform selection control unit 60 controls the inverse orthogonal transform unit 52 and the filter processing unit 54 to perform the filtering and the type of orthogonal transform according to a predetermined method (step S26).

[0161] As described above, the inverse orthogonal transform selection control unit 60 determines the type of inverse orthogonal transform process to be applied and whether or not to perform filtering by referring to the transformation type identification signal obtained from the image encoding device 1, and controls the inverse orthogonal transform unit 52 and the filtering unit 54. As a result, the residual component in the predicted residual signal can be reduced on the image encoding device 1 side, decoding can be performed with less information, and the encoding efficiency can be improved in a series of processes on the encoding and decoding sides.

[0162] In particular, even when predicting without referring to adjacent pixels, the encoding device 1 uses an inverse orthogonal transform process (e.g., IDST) with closed ends for the transform basis to apply to the block of predicted residual signal obtained based on the block of the predicted image after the filtering process. This matches not only the phase but also the characteristics of the transform basis corresponding to the endpoints of the prediction region, thereby significantly improving the encoding efficiency.

[0163] With the image encoding device 1 and image decoding device 5 configured as described above, the signal error that occurs between the predicted signal of the predicted image and the adjacent decoded block signal is reduced, improving encoding efficiency. Therefore, it is possible to realize an image encoding device and image decoding device for a video encoding method with high encoding efficiency.

[0164] [Fourth Embodiment] (Image encoding device) Next, the filter processing unit 12 related to the predicted image, which is a major component of the present invention in the image coding device 1 of the fourth embodiment of the present invention, will be described with reference to Figures 19 and 20, and a specific typical example of the image coding device 1 will be described with reference to Figures 4, 21 and 22. The following description will focus on the differences from the embodiments described above. That is, in this embodiment, the same reference numerals are used for components that are the same as in the embodiments described above, and further detailed descriptions thereof will be omitted.

[0165] In general, the image coding device 1 of this embodiment applies, for example, a low-pass filter to the prediction image used for inter-prediction, similar to intra-prediction, by utilizing the encoded and decoded adjacent signals. This reduces the prediction residual signal for the leftmost and uppermost regions of the prediction image used for inter-prediction, and divides the block of prediction residual signal into blocks, applying either DST or DCT depending on the direction of the low-pass filter. Note that under the current H.265 standard, the block size of the prediction image to which DST can be applied is limited to small ones (e.g., 4x4). Therefore, in the image coding device 1 of this embodiment, for inter-prediction in larger prediction image blocks, after applying the low-pass filter to the prediction image, the block is divided into a predetermined block size (e.g., the block size to which DST can be applied under the standard), and then it is determined without a flag whether to apply DST or DCT, and the corresponding orthogonal transformation is executed.

[0166] Figure 19(a) is a block diagram of the area around the filter processing unit 12 related to the predicted image in the image encoding device 1 of this embodiment, and Figure 19(b) is an explanatory diagram showing an example of filtering the predicted image by the filter processing unit 12.

[0167] As shown in Figure 19(a), the image encoding device 1 of this embodiment is configured to include an adjacent pixel non-reference prediction unit 11, a filter processing unit 12, a prediction residual signal generation unit 13, and an orthogonal transformation unit 14. The orthogonal transformation unit 14 includes a block division unit 141 and an orthogonal transformation selection application unit 142.

[0168] In this embodiment, the filter processing unit 12 is a functional unit that generates a new predicted image consisting of a new predicted signal by applying a low-pass filter to the predicted signal using the decoded signals (decoded adjacent signals) that are adjacent to the left and above the predicted image, respectively, from the pixel signals (predicted signals) within the block of the predicted image generated by the adjacent pixel non-reference prediction unit 11, and outputs this to the predicted residual signal generation unit 13.

[0169] For example, as shown in Figure 19(b), if a block Bkp of a predicted image consisting of 8x8 pixel signals p corresponding to block Bko of the encoding region of the original image is generated by the adjacent pixel non-reference prediction unit 11, the filter processing unit 12 uses the decoded adjacent signal So around the block boundary of the block Bkp of the predicted image adjacent to the predicted image, and uses the leftmost and uppermost pixel signals (predicted signals) of the predicted image as the signal Sf of the region to be filtered, and applies low-pass filtering such as a smoothing filter (double arrows shown in the figure), such as a horizontal filter, a vertical filter, and a square filter, respectively, to generate a new predicted image consisting of the predicted signals.

[0170] More specifically, in the example shown in Figure 19(b), for the upper smoothing filter target pixels of block Bkp in the predicted image, smoothing filtering is performed using decoded adjacent signals (pixel signals of the locally decoded image) located at the same coordinates horizontally, and for the left smoothing filter target pixels, smoothing filtering is performed using decoded adjacent signals (pixel signals of the locally decoded image) located at the same coordinates vertically. For target pixels located both above and to the left, smoothing filtering is performed using pixels from the upper and left locally decoded images. Note that in the example shown in Figure 19(b), a 3-tap smoothing filter using 2 reference pixels is applied. For example, a low-pass filter such as 1 / 4

[0121] is applied. Note that the number of taps and reference pixels for the low-pass filter are not limited to this example.

[0171] The orthogonal transformation unit 14 includes a block division unit 141 and an orthogonal transformation selection and application unit 142. The block division unit 141 divides the blocks of the predicted residual signal input from the predicted residual signal generation unit 13 into predetermined block shapes and outputs them to the orthogonal transformation selection and application unit 142. The orthogonal transformation selection and application unit 142 selects and applies to each block of the divided predicted residual signal a combination of vertical / horizontal DST and vertical / horizontal DCT, depending on the position where the filter processing by the filter processing unit 12 has been applied.

[0172] For example, as shown in Figure 19(b), the block division unit 141 divides a block of predicted residual signals for a block Bkp of a predicted image consisting of 8x8 pixel signals p into four blocks of 4x4 predicted residual signals (upper left, upper right, lower left, and lower right) and outputs them to the orthogonal transformation selection application unit 142. At this time, the orthogonal transformation selection application unit 142 applies multiple types of orthogonal transformations (combinations of vertical / horizontal DST and vertical / horizontal DCT) according to the position to which the smoothing filter processing was applied to the divided blocks of predicted residual signals located at the upper and left ends, including the pixel positions to which the smoothing filter processing was applied, so that it can handle transformation coefficients that utilize the high correlation with the decoded adjacent signal blocks. For the other divided blocks, the vertical / horizontal DCT orthogonal transformation processing is applied. Note that vertical direction means the vertical direction, and horizontal direction means the horizontal direction. Here, the orthogonal transformation process consisting of vertical DST and horizontal DCT is defined as the first orthogonal transformation process, the orthogonal transformation process consisting of horizontal DST and vertical DCT is defined as the second orthogonal transformation process, the orthogonal transformation process consisting of vertical and horizontal DST is defined as the third orthogonal transformation process, and the orthogonal transformation process consisting of vertical and horizontal DCT is defined as the fourth orthogonal transformation process.

[0173] In the example shown in Figure 19(b), a block of predicted residual signal corresponding to a block Bkp of a predicted image consisting of 8x8 pixel signals p is divided into four blocks of 4x4 predicted residual signal (upper left, upper right, lower left, and lower right). However, for blocks of the predicted residual signal corresponding to blocks of the predicted image with a block size larger than twice the size obtained by doubling the vertical and horizontal dimensions (8x8 in this example) of a predetermined block size (4x4 in this example), it is preferable to set the division blocks adjacent to the decoded adjacent signals (the division blocks located at the top and left ends in this example) to the predetermined block size (4x4 in this example), and to divide the other blocks to the largest possible block size, and then appropriately subdivide them according to the characteristics of the image to be encoded. For example, depending on the characteristics of the image to be encoded, as shown in Figure 20(a), the division blocks located at the top and left ends of a 32x32 block size predicted residual signal block can be set to 4x4, and the division blocks can be made into increasingly larger block sizes as they move away from the top and left ends. By performing block division in this manner, the subsequent orthogonal transformation selection and application unit 142 can select and apply an orthogonal transformation suitable for the signal characteristics of the divided blocks of the predicted residual signal that include the pixel positions to be processed by the smoothing filter.

[0174] For example, as shown in Figure 20(a), the orthogonal transformation selection application unit 142 applies a first orthogonal transformation process of vertical DST and horizontal DCT to the uppermost division block of the predicted residual signal Bkt1 if the pixel to which the smoothing filter processing has been applied is at the top (i.e., the direction of application of the smoothing filter processing is vertical) (see Figure 20(b)). If the pixel to which the smoothing filter processing has been applied is at the leftmost position (i.e., the direction of application of the smoothing filter processing is horizontal), the second orthogonal transformation process of horizontal DST and vertical DCT is applied to the leftmost division block of the predicted residual signal Bkt2 (see Figure 20(c)). Then, if the pixel to which the smoothing filter processing has been applied is at the top and also at the leftmost position (i.e., the direction of application of the smoothing filter processing is the direction of application of the angular filter processing), the third orthogonal transformation process of vertical and horizontal DST is applied to the division block Bkt3 located in the angular region. Furthermore, for divided blocks Bkt4 of the predicted residual signal where no pixels have been subjected to smoothing filtering, the fourth orthogonal transform of the DCT is applied in the vertical and horizontal directions. The same orthogonal transform is performed when the block is divided as shown in Figure 19(b). This makes it possible to handle transformation coefficients that make the most of the high correlation of the decoded adjacent signals with respect to the block.

[0175] Next, as a typical example of the adjacent pixel non-reference prediction unit 11, an inter-prediction unit 11a using inter-prediction will be described, and the configuration of the image encoding device 1 and an example of its operation, in which the prediction signal of the predicted image block is filtered by the filter processing unit 12, will be described with reference to Figures 21 to 23.

[0176] Figure 21 is a block diagram showing one embodiment of the image coding device 1 of this embodiment. The image coding device 1 shown in Figure 21 comprises a preprocessing unit 10, an interpretation unit 11a, a filter processing unit 12, a prediction residual signal generation unit 13, an orthogonal transformation unit 14, a quantization unit 15, an inverse quantization unit 16, an inverse orthogonal transformation unit 17, a decoded image generation unit 18, an in-loop filter unit 19, a frame memory 20, an intraprediction unit 21, a motion vector calculation unit 22, an entropy coding unit 23, and a prediction image selection unit 24.

[0177] The orthogonal transformation unit 14 applies a predetermined orthogonal transformation process to the predicted residual signal input from the predicted residual signal generation unit 13 and generates a signal of the transformation coefficients. In particular, the orthogonal transformation unit 14 includes a block division unit 141 and an orthogonal transformation selection application unit 142. The block division unit 141 divides the blocks of the predicted residual signal input from the predicted residual signal generation unit 13 into predetermined block shapes and outputs them to the orthogonal transformation selection application unit 142. The orthogonal transformation selection application unit 142 applies multiple types of orthogonal transformation processes (a combination of vertical / horizontal DST and vertical / horizontal DCT) to each divided block of the divided predicted residual signal, according to the position to which the filtering process by the filtering processing unit 12 was applied. The block division parameters related to these block division processes are output to the entropy coding unit 23 (and the inverse orthogonal transformation unit 17) as one of the coding parameters that are supplementary information for the coding process.

[0178] The inverse orthogonal transform unit 17 reconstructs the predicted residual signal by applying an inverse orthogonal transform process to the signal of the conversion coefficients after inverse quantization obtained from the inverse quantization unit 16, and outputs it to the decoded image generation unit 18. In particular, the inverse orthogonal transform unit 17 determines the type of orthogonal transform process for the block-divided conversion coefficients, as in Figure 20, applies the corresponding inverse orthogonal transform process to each orthogonal transform process, combines each divided block of the predicted residual signal obtained in parallel processing, reconstructs a block corresponding to the block size of the predicted image, and outputs it to the decoded image generation unit 18.

[0179] An example of the operation of the filter processing unit 12 will be explained with reference to Figure 22. As shown in Figure 22, when the filter processing unit 12 receives a predicted image from the intra prediction unit 21 or the inter prediction unit 11a (step S1), it identifies whether the predicted image is an adjacent pixel non-reference prediction, that is, in this example, whether it is an inter prediction based on signal selection by the predicted image selection unit 24 (step S2). The same applies when the adjacent pixel non-reference prediction is the intra block copy prediction or cross component signal prediction described above.

[0180] Next, if the filter processing unit 12 determines that the predicted image is not an adjacent pixel non-reference prediction, i.e., an intra-prediction in this example (step S2: No), it determines whether or not to perform filtering according to a predetermined processing method (step S5). If filtering is not performed on the predicted image of an intra-prediction (step S5: No), the filter processing unit 12 outputs the predicted residual signal to the prediction residual signal generation unit 13 without performing filtering on the predicted signal (step S6). On the other hand, if filtering is performed on the predicted image of an intra-prediction (step S5: Yes), the filter processing unit 12 proceeds to step S3. Note that filtering on the predicted image of an intra-prediction can be done in the same way as currently defined H.265 and is not directly related to the essence of the present invention, so the explanation of step S3 will describe an example of performing filtering on the predicted image of an inter-prediction.

[0181] If the filter processing unit 12 determines that the predicted image is an adjacent pixel non-reference prediction, that is, in this example, an inter-prediction (Step S2: Yes), it selects the pixel signals (prediction signals) of a predetermined filtering target region from the predicted image (Step S3). For example, in the example shown in Figure 19(b), the leftmost and uppermost pixel signals (prediction signals) of the predicted image are defined as the filtering target region Sf.

[0182] Next, the filter processing unit 12 uses the decoded signals adjacent to the left and above the predicted image (decoded adjacent signals) to apply a low-pass filter to the pixel signals (predicted signals) of the selected filtering target region, thereby generating a new predicted image consisting of new predicted signals, and outputs it to the predicted residual signal generation unit 13 (step S4).

[0183] Furthermore, the filtering performed by the filter processing unit 12 can be any low-pass filter, and filtering other than smoothing filters can be applied. The region of the decoded adjacent signal and the weight coefficients used for filtering are not limited to the example shown in Figure 19(b). In addition, the filter type, the region to be filtered, and the weight coefficients related to the filtering performed by the filter processing unit 12 can be configured to be transmitted as supplementary information to the encoding process, but they do not need to be transmitted if they are predetermined between the sender and receiver (between encoding and decoding).

[0184] By applying a filter to the predicted image using the filter processing unit 12, regardless of the type of orthogonal transformation processing performed by the orthogonal transformation unit 14, the signal error between the leftmost and uppermost pixel signals (predicted signals) in the predicted image and the adjacent decoded signals is reduced, thereby improving the encoding efficiency of the video.

[0185] Next, an example of the operation of the orthogonal transformation unit 14 will be explained with reference to Figure 23. As shown in Figure 23, the orthogonal transformation unit 14 divides the blocks of predicted residual signals input from the predicted residual signal generation unit 13 into predetermined block shapes using the block division unit 141 (step S11). At this time, for blocks of predicted images with a block size larger than 8 × 8 (for example, for a predicted image with a block size of 32 × 32), the block division unit 141 sets the division blocks of the predicted residual signals located at the top and left ends to a predetermined block size (4 × 4 in this example), and divides the other blocks to the largest possible block size, and then appropriately performs further division according to the characteristics of the image to be encoded.

[0186] Next, the orthogonal transformation unit 14, when applying orthogonal transformation processing to the divided blocks of the divided predicted residual signals by the orthogonal transformation selection application unit 142, determines whether or not the divided block includes a pixel position to which filtering by the filtering unit 12 has been applied (i.e., whether or not it is the predicted residual signal of a block that includes the uppermost or leftmost pixel position of the predicted image) (step S12). If it is determined that the divided block does not include a pixel position to which filtering has been applied (step S12: No), it applies the fourth orthogonal transformation processing of the DCT in the vertical and horizontal directions to the divided block of the predicted residual signal (step S15).

[0187] On the other hand, when it is determined that the divided block includes a pixel position to which filtering by the filtering processing unit 12 has been applied (step S12: Yes), it is determined (step S13) whether only vertical filtering by the filtering processing unit 12 has been applied (i.e., whether it is a predicted residual signal of a block that includes only the uppermost pixels of the predicted image), and if only vertical filtering has been applied (step S13: Yes), the first orthogonal transformation processing of vertical DST and horizontal DCT is applied to the divided block of the predicted residual signal (step S16).

[0188] Next, if the vertical filtering has not been applied (step S12: No), it is determined whether only the horizontal filtering by the filtering unit 12 has been applied (i.e., whether it is a predicted residual signal of a block containing only the leftmost pixel of the predicted image) (step S14). If only the horizontal filtering has been applied (step S14: Yes), the second orthogonal transformation processing of horizontal DST and vertical DCT is applied to the divided block of the predicted residual signal (step S17).

[0189] Then, when it is determined that horizontal and vertical filtering has been applied by the filtering processing unit 12 (i.e., when it is the predicted residual signal of a block that includes the uppermost and leftmost pixel position of the predicted image) (step S14: No), a third orthogonal transformation process of the vertical and horizontal DST is applied to the divided block of the predicted residual signal (step S18).

[0190] In Figure 23, the process is explained as sequential orthogonal transformation for better understanding, but in practice, parallel processing is possible. This allows for the use of transformation coefficients that make the most of the high correlation between the decoded adjacent signals and the blocks.

[0191] (Image decoding device) Next, the filter processing unit 54 and its peripheral function blocks, which are the main components of the image decoding device 5 according to the present invention, will be described with reference to Figure 24, and a specific typical example of the image decoding device 5 will be described with reference to Figure 25.

[0192] Figure 24 is a block diagram of the area around the filter processing unit 54 related to the predicted image in the image decoding device 5 of this embodiment. As shown in Figure 24, the image decoding device 5 of this embodiment is configured to include an inverse orthogonal transform unit 52, an adjacent pixel non-reference prediction unit 53, a filter processing unit 54, and a decoded image generation unit 55. The inverse orthogonal transform unit 52 also includes an inverse orthogonal transform selection application unit 521 and a block reconstruction unit 522.

[0193] The inverse orthogonal transform unit 52 receives the stream signal transmitted from the image encoding device 1, performs an inverse orthogonal transform on the transformation coefficients restored through entropy decoding and inverse quantization processing, and outputs the resulting predicted residual signal to the decoded image generation unit 55. In particular, the inverse orthogonal transform unit 522 includes an inverse orthogonal transform selection application unit 521 and a block reconstruction unit 522. The inverse orthogonal transform selection application unit 521 determines the type of orthogonal transform processing for the transformation coefficients that have been divided into blocks by the image encoding device 1, as shown in Figure 23, and performs an inverse orthogonal transform processing corresponding to the orthogonal transform processing in the image encoding device 1, outputting the divided blocks of the predicted residual signal to the block reconstruction unit 522. The block reconstruction unit 522 combines the divided blocks of the predicted residual signal obtained in parallel processing to reconstruct blocks corresponding to the block size of the predicted image, and outputs them to the decoded image generation unit 55.

[0194] Next, as a typical example of the adjacent pixel non-reference prediction unit 53, we will describe the inter-prediction unit 53a using inter-prediction, and then, with reference to Figure 25, we will describe the configuration of the image decoding device 5 and an example of its operation in which the predicted signal from the inter-prediction is filtered by the filter processing unit 54.

[0195] Figure 25 is a block diagram showing one embodiment of the image decoding device 5 of this embodiment. The image decoding device 5 shown in Figure 25 comprises an entropy decoding unit 50, an inverse quantization unit 51, an inverse orthogonal transformation unit 52, an interpretation unit 53a, a filter processing unit 54, a decoded image generation unit 55, an in-loop filter unit 56, a frame memory 57, an intraprediction unit 58, and a predicted image selection unit 59.

[0196] The inverse orthogonal transform unit 52 reconstructs the predicted residual signal by applying an inverse orthogonal transform process corresponding to the orthogonal transform process of the image coding device 1 to the transformation coefficients obtained from the inverse quantization unit 51, and outputs it to the decoded image generation unit 55. In particular, the inverse orthogonal transform unit 522 includes an inverse orthogonal transform selection application unit 521 and a block reconstruction unit 522. The inverse orthogonal transform selection application unit 521 applies an inverse orthogonal transform process to the transformation coefficients that have been divided into blocks by the image coding device 1, and outputs the divided blocks of the predicted residual signal to the block reconstruction unit 522. The block reconstruction unit 522 combines the divided blocks of the predicted residual signal, which can be processed in parallel, to form a block corresponding to the block size of the predicted image, and outputs it to the decoded image generation unit 55.

[0197] The filter processing unit 54 operates in the same manner as illustrated with reference to Figure 22. Specifically, when the predicted image is an adjacent pixel non-reference prediction, i.e., an inter-prediction in this example, the filter processing in the filter processing unit 54 selects a pixel signal (prediction signal) from a predetermined filtering target area of ​​the predicted image, and uses the decoded adjacent signal (decoded adjacent signal) adjacent to the left or above the predicted image to apply a low-pass filter to the selected pixel signal (prediction signal) of the filtering target area to generate a new predicted image consisting of a new predicted signal, which is then output to the decoded image generation unit 55. For example, as illustrated in Figure 19(b), the filter processing unit 54 uses the decoded adjacent signal So around the block boundary of block Bkp of the predicted image to apply a smoothing filter with predetermined weight coefficients to the horizontal filter area, vertical filter area, and angular filter area of ​​the filtering target area Sf within the pixel signal (prediction signal) of block Bkp of the predicted image, and outputs it to the decoded image generation unit 55.

[0198] With the image encoding device 1 and image decoding device 5 configured as described above, the signal error occurring between the predicted signal of the predicted image and the adjacent decoded block signal is reduced, improving encoding efficiency. Therefore, an image encoding device and image decoding device for a video encoding method with high encoding efficiency can be realized. In other words, the residual component in the predicted residual signal can be reduced, and encoding efficiency can be improved by reducing the amount of information to be encoded and transmitted.

[0199] [Fifth Embodiment] Next, the configurations of the filter processing units 12 and 54 in the image encoding device 1 and image decoding device 5 of the fifth embodiment of the present invention will be described with reference to Figure 26. In the filter processing units 12 and 54 of the fourth embodiment described above, an example was described in which a new predicted image is generated by applying a low-pass filter to the predicted signal using only the "decoded signals (decoded adjacent signals)" that are adjacent to the left and above the predicted image, respectively, from the pixel signals (predicted signals) within the block of the predicted image generated by adjacent pixel non-reference prediction.

[0200] On the other hand, the filter processing units 12 and 54 of this embodiment acquire a locally decoded image for each divided block based on the block division process of the block division unit 141. Each time this locally decoded image is acquired, the pixel signals (predicted signals) within the blocks of the predicted image generated by adjacent pixel non-reference prediction are replaced. Then, in units of the block size of the divided blocks of the predicted residual signal, a low-pass filter is applied according to the respective block size using the decoded signals (decoded adjacent signals) that are adjacent to the left and above the predicted image from among the replaced pixel signals (predicted signals) within the blocks of the predicted image, and the locally decoded image for each divided block.

[0201] Therefore, as shown in Figures 26(a) and (b), the filter processing units 12 and 54 acquire a local decoded image for each divided block based on the block division process of the block division unit 141, and each time this local decoded image is acquired, they replace the pixel signals (predicted signals) in the block of the predicted image generated by adjacent pixel non-reference prediction to update the predicted image, and apply a low-pass filter process according to the respective block size to this updated predicted image using the decoded signals adjacent to the left and above (decoded adjacent signals) and the local decoded image for each divided block. In this respect, they differ from the configurations shown in Figures 19(a) and 24, respectively, but the operation of the other components is the same.

[0202] As an example of its operation, the block reconstruction unit 522 in the inverse orthogonal transform unit 52 shown in Figure 26(b) differs from the configuration shown in Figure 24 in that when determining the position of the divided blocks to be processed and reconstructing them to the block size corresponding to the predicted image, it interpolates areas other than the divided blocks to be processed with dummy data. This makes it possible to configure the system so that a local decoded image can be obtained for each divided block. Alternatively, a separate loop may be configured so that a local decoded image can be obtained for each divided block. Therefore, the inverse orthogonal transform unit 17 in the image encoding device 1 of this embodiment can be configured in the same way as the inverse orthogonal transform unit 52 of this embodiment. In addition, as is also the case in the fourth embodiment, the frame memory 20 can store the local decoded image for each divided block so that it can be used as a reference signal.

[0203] Figure 27 is a flowchart relating to the filtering processes 12 and 54 of the predicted image in the image encoding device 1 or image decoding device 5 of this embodiment. Figure 28 is an explanatory diagram relating to the filtering processes 12 and 54 of the predicted image in the image encoding device 1 or image decoding device 5 of this embodiment.

[0204] First, when the filter processing units 12 and 54 receive a predicted image as input, they acquire a local decoded image for each divided block based on the block division process of the block division unit 141, and replace it with the predicted image to generate an updated predicted image (step S21). For example, as illustrated in Figure 28, the block Bkp of the 8x8 predicted image is sequentially replaced and updated with the block BkL of the local decoded image, which is divided into four blocks in the order shown by the white arrows. The following process shows the operation when the predicted image is updated with the local decoded image for each divided block (for example, "STEP 1" in the figure).

[0205] Next, the filter processing units 12 and 54 identify whether the predicted image is an adjacent pixel non-reference prediction, for example, whether it is an inter-prediction (step S22). The same applies if the adjacent pixel non-reference prediction is the intra-block copy prediction or cross-component signal prediction mentioned above.

[0206] Next, if the filter processing units 12 and 54 determine that the predicted image is not an adjacent pixel non-reference prediction, for example, an intra-prediction (step S22: No), they determine whether or not to perform filtering according to a predetermined processing method (step S25). If filtering is not performed on the predicted image of an intra-prediction (step S25: No), the filter processing units 12 and 54 output the predicted residual signal generation unit 13 and the decoded image generation unit 55 without performing filtering on the predicted signal (step S26). On the other hand, if filtering is performed on the predicted image of an intra-prediction (step S25: Yes), the filter processing units 12 and 54 proceed to step S23. Note that filtering on the predicted image of an intra-prediction can be done in the same way as currently defined H.265 and is not directly related to the essence of the present invention, so the explanation of step S23 will describe an example of performing filtering on the predicted image of an inter-prediction.

[0207] If the filter processing units 12 and 54 determine that the predicted image is an adjacent pixel non-reference prediction, that is, in this example, an inter-prediction (step S22: Yes), they select the pixel signals (predicted signals) of a predetermined filtering target region from the predicted image (step S23). For example, as illustrated in Figure 28, in this example, the leftmost and uppermost pixel signals (predicted signals) for adjacent signals are defined as the filtering target region Sf in 4x4 divided block units.

[0208] Next, the filter processing units 12 and 54 apply a low-pass filter to the pixel signals (predicted signals) of the selected filtering target region using the decoded signals (decoded adjacent signals) adjacent to the left and / or above each divided block and the local decoded image of the predicted image replaced with the locally decoded image of the divided block, thereby generating a new predicted image consisting of a new predicted signal, which is output to the predicted residual signal generation unit 13 and the decoded image generation unit 55 (step S24).

[0209] By applying a low-pass filter using the decoded signals adjacent to the left and / or top of the predicted image (decoded adjacent signals) and the locally decoded images adjacent to the left and / or top of each divided block, for example, as illustrated in Figure 28, the pixel signals of the 8x8 predicted image block Bkp are smoothed not only on its top and left sides, but also at the boundary positions of the divided blocks which form a cross shape in this example (for example, shown as "STEP 2").

[0210] By configuring the filter processing units 12 and 54 as in this embodiment and applying them to the image encoding device 1 and the image decoding device 5, respectively, it is possible to improve encoding efficiency while suppressing the resulting degradation of image quality. In particular, even for the signals of division blocks other than the division block located on the leftmost or topmost side of the predicted image, which can result in large predicted residual signals (blocks located inside the predicted image when the image is divided into blocks), the values ​​of the predicted residual signals on the leftmost and topmost sides of each division block can be reduced.

[0211] In this embodiment, since all divided blocks include the pixel positions to be filtered, vertical or horizontal DST is always applied. Therefore, in this embodiment, when dividing the blocks, it is preferable to divide the divided blocks located at the top and left ends of the corresponding predicted residual signal block to a specified block size (4x4 in this example), as in the fourth embodiment, and divide the other blocks to the largest possible block size, and then appropriately subdivide them according to the characteristics of the image to be encoded. It is not necessary to divide the blocks so that the block size increases as you move away from the top and left ends; all blocks can be divided to a specified size (4x4 in this example) to which DST can be applied, and DST can be applied according to the processing order shown in Figure 23.

[0212] Furthermore, comparing the fourth embodiment described above with this embodiment, the fourth embodiment has the advantage of enabling parallel processing because each divided block of the predicted residual signal can be processed independently, while this embodiment can sequentially utilize the locally decoded images of each divided block, thus allowing for a greater improvement in coding efficiency.

[0213] Furthermore, since the block shape of the block division can be set variably for each block to be encoded, and the block division parameter can be transmitted as one of the encoding parameters, an embodiment combining the fourth embodiment and this embodiment can also be created.

[0214] Each embodiment of the image encoding device 1 and image decoding device 5 can function as a computer, and the program for implementing each component according to the present invention is stored in memory located inside or outside the computer. The central processing unit (CPU) or the like in the computer can control the program, which describes the processing content for implementing the functions of each component, by reading it from memory as appropriate, thereby enabling the computer to implement the functions of each component of the image encoding device 1 and image decoding device 5 in each embodiment. Here, the functions of each component may be implemented in part of the hardware.

[0215] Although the present invention has been described above with reference to specific embodiments, the present invention is not limited to the examples described above and can be modified in various ways without departing from the technical concept. For example, in the examples of each embodiment described above, the examples described were in which the encoding target area, the predicted image, and the block size of the orthogonal transformation target were the same, but the same can be applied when the block size of the predicted image is smaller than that of the encoding target area, or when the block size of the orthogonal transformation target is smaller than that of the predicted image. Furthermore, the orthogonal transformation process can be the same as that used in the encoding process, depending on the application, or a different one may be used.

[0216] Furthermore, in the examples of the embodiments described above, the image decoding device 5 was shown to decode conversion coefficients related to the predicted residual signal encoded based on the predicted image filtered by the corresponding image encoding device 1. However, for applications where the accuracy of restoring block edges to the original image is not a concern, the image decoding device 5 according to the present invention can decode signals encoded without filtering using the same process. [Industrial applicability]

[0217] According to the present invention, it is possible to realize an image encoding device and an image decoding device for a video encoding method with high encoding efficiency, making them useful for applications where it is desirable to improve the encoding efficiency of video transmission. [Explanation of Symbols]

[0218] 1 Image encoding device 5 Image Decoder 10 Pre-processing 11. Adjacent Pixel Non-Reference Prediction Unit 11a Interpretation Unit 12 Filtering section 13 Predicted residual signal generation unit 14. Orthogonal Transformation Section 15 Quantization section 16 Inverse quantization section 17 Inverse orthogonal transformation section 18 Decoded Image Generation Unit 19. In-loop filter section 20 frames memory 21 Intra Prediction Unit 22 Motion Vector Calculation Unit 23 Entropy coding unit 24 Predictive Image Selection Section 25 Orthogonal Transformation Selection Control Unit 50 Entropy Decoder 51 Inverse quantization section 52 Inverse orthogonal transformation section 53 Adjacent Pixel Non-Reference Prediction Unit 53a Interpretation Unit 54 Filtering section 55 Decoded Image Generation Unit 56 Loop-based filter section 57 Frame Memory 58 Intra Prediction Unit 59 Predictive Image Selection Section 60 Inverse orthogonal transformation selection control unit 101 Horizontal correlation determination unit 102 Vertical Correlation Determination Unit 103 Filter Processing Determination Unit 104 Filter Processing Execution Unit 141 Block division section 142 Orthogonal Transformation Selection Application Unit 521 Inverse orthogonal transformation selection and application unit 522 Block Reconstruction Unit

Claims

1. An image decoding device that decodes signals encoded by dividing frames constituting a moving image into blocks, A prediction unit generates a block-unit predicted image containing the predicted signal by performing signal prediction for each pixel signal in block units, An acquisition unit that acquires a control identification signal indicating the type of selection control selected by the encoding side, A filter processing unit generates a new block of predicted image by applying a filter to the predicted signal using a decoded adjacent signal adjacent to the block of predicted image based on the control identification signal, An inverse quantization unit that reconstructs the converted coefficient signal by applying a corresponding inverse quantization process to the quantized converted coefficients, The system includes an inverse transformer that, based on the control identification signal, determines whether or not to apply a specific inverse transformer, including inverse DST and inverse DCT, to the conversion coefficient signal, and if it is determined that the specific inverse transformer should be applied, reconstructs the block-level predicted residual signal by applying the specific inverse transformer. Image decoding device.

2. An image decoding method that decodes a signal encoded by dividing the frames constituting a moving image into blocks, The process involves generating a block-based predicted image containing the predicted signal by performing signal prediction for each pixel signal in a block-based manner, and The steps include obtaining a control identification signal indicating the type of selection control selected by the encoding side, The steps include generating a new block of the predicted image by filtering the predicted signals located at the boundaries of the block of the predicted image using the decoded adjacent signals adjacent to the block of the predicted image, The steps include:

1. Reconstructing the transformed coefficient signal by applying the corresponding inverse quantization process to the quantized transformed coefficients; The process includes the steps of: determining whether to apply a specific inverse transformation process, including inverse DST and inverse DCT, to the conversion coefficient signal based on the control identification signal; and, if it is determined that the specific inverse transformation process should be applied, reconstructing the block-level predicted residual signal by applying the specific inverse transformation process. Image decoding method.

Citation Information

Patent Citations

  • Method, device and program for generating prediction signal

    JP2011066569A

  • Encoding apparatus, decoding apparatus, and program

    JP2011205602A

  • Image encoder, image decoder and program

    JP2014036278A

  • Apparatus and method for encoding and decoding using alternative converter accoding to the correlation of residual signal

    US20090238271A1

  • Methods and apparatus for constrained transforms for video coding and decoding having transform selection

    WO2011112239A1