Image decoding device, image decoding method, and program
The image decoding apparatus addresses inefficiencies in coding efficiency and processing time by using a decoding unit and intra-frame prediction units to strategically evaluate and prioritize block vector coordinates in IntraTMP, resulting in improved performance.
Patent Information
- Application Number
- JP2023201282
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-29
- Publication Date
- 2025-06-10
AI Technical Summary
Existing image decoding technologies face inefficiencies in coding efficiency due to narrow search ranges for block vectors (BV) in Intra Template Matching Prediction (IntraTMP), leading to increased processing times for both encoding and decoding.
The proposed image decoding apparatus and method include a decoding unit that decodes control information and quantization values, an inverse quantization unit, an inverse transform unit, and intra-frame prediction units that generate predicted pixels based on decoded pixels and control information. The second intra-frame prediction unit preferentially evaluates coordinates where a block vector is likely to exist, thereby generating a second predicted pixel and accumulating decoded pixels for inter-frame prediction.
This approach enhances coding efficiency while reducing processing times by strategically evaluating and prioritizing block vector coordinates, thus improving the overall performance of image decoding.
Smart Images

Figure 2025086972000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image decoding apparatus, an image decoding method, and a program.
Background Art
[0002] Non-Patent Documents 1 to 3 disclose Intra Template Matching Prediction (IntraTMP).
[0003] IntraTMP refers to pixels that match or are similar by template matching from the decoded pixel region of the frame to be decoded, and uses them as predicted pixels for the block to be decoded.
[0004] Specifically, the image encoding apparatus uses the decoded neighboring pixels as a template, searches for coordinates with a small template matching cost from the same frame, and sets the displacement amount to such coordinates as a block vector (hereinafter referred to as BV).
[0005] The image encoding apparatus constructs a BV list (reference list) in ascending order of the template matching cost and encodes the index of such a BV list.
[0006] The image decoding apparatus also searches in the same way and decodes the BV from the index of such a BV list by reconstructing the above BV list.
[0007] As shown in FIG. 2, the image decoding apparatus uses the pixels copied from the reference block indicated by the BV of the block to be decoded as predicted pixels.
Prior Art Documents
Non-Patent Documents
[0008]
Non-Patent Document 1
[0009] In Non-Patent Documents 1 and 2, since the search range of BV is narrow, there is a problem that the coding efficiency is not sufficient.
[0010] In Non-Patent Document 3, by expanding the search range of BV for IntraTMP, some of the problems in Non-Patent Documents 1 and 2 described above are solved.
[0011] However, there is a problem that both the encoding processing time and the decoding processing time increase when the search range of BV is expanded. In particular, when further expanding the search range of BV proposed by Non-Patent Document 3, such a problem becomes more prominent.
[0012] Therefore, the present invention has been made in view of the above problems, and an object thereof is to provide an image decoding apparatus, an image decoding method, and a program with high coding efficiency.
Means for Solving the Problem
[0013] A first feature of the present invention is an image decoding apparatus, which includes a decoding unit that decodes control information and quantization values, an inverse quantization unit that inverse quantizes the quantization values to obtain transform coefficients, an inverse transform unit that inverse transforms the transform coefficients to obtain prediction residuals, a first in-frame prediction unit that generates a first predicted pixel based on the decoded pixels and the control information, a second in-frame prediction unit that preferentially evaluates coordinates where a block vector is likely to exist based on the decoded pixels and the control information, and generates a second predicted pixel based on the evaluation result, an accumulation unit that accumulates the decoded pixels, an inter-frame prediction unit that generates a third predicted pixel based on the accumulated decoded pixels and the control information, and an adder that adds the prediction residuals and the first to third predicted pixels to obtain the decoded pixels.
[0014] A second feature of the present invention is an image decoding method, which includes a step of decoding control information and quantization values, a step of inverse quantizing the quantization values to obtain transform coefficients, a step of inverse transforming the transform coefficients to obtain prediction residuals, a step of generating a first predicted pixel based on the decoded pixels and the control information, a step of preferentially evaluating coordinates where a block vector is likely to exist based on the decoded pixels and the control information, and generating a second predicted pixel based on the evaluation result, a step of accumulating the decoded pixels, a step of generating a third predicted pixel based on the accumulated decoded pixels and the control information, and a step of adding the prediction residuals and the first to third predicted pixels to obtain the decoded pixels.
[0015] A third feature of the present invention is a program that causes a computer to function as an image decoding device. The image decoding device includes a decoding unit that decodes control information and quantization values, an inverse quantization unit that inverse quantizes the quantization values to obtain transform coefficients, an inverse transform unit that inverse transforms the transform coefficients to obtain prediction residuals, a first intra-frame prediction unit that generates a first predicted pixel based on the decoded pixels and the control information, a second intra-frame prediction unit that preferentially evaluates coordinates where a block vector is likely to exist based on the decoded pixels and the control information, and generates a second predicted pixel based on the evaluation result, an accumulation unit that accumulates the decoded pixels, an inter-frame prediction unit that generates a third predicted pixel based on the accumulated decoded pixels and the control information, and an adder that adds the prediction residuals and the first to third predicted pixels to obtain the decoded pixels.
Advantages of the Invention
[0016] According to the present invention, it is possible to provide an image decoding device, an image decoding method, and a program with high coding efficiency.
Brief Description of the Drawings
[0017]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
[0018] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Note that the components in the following embodiments can be appropriately replaced with existing components and the like, and various variations including combinations with other existing components are possible. Therefore, the description of the following embodiments does not limit the content of the invention described in the claims.
[0019] <First Embodiment> Hereinafter, with reference to FIGS. 1 to 10, the image decoding apparatus 200 according to the present embodiment will be described. FIG. 1 is a diagram showing an example of the functional blocks of the image decoding apparatus 200 according to the present embodiment.
[0020] As shown in FIG. 1, the image decoding apparatus 200 includes a code input unit 210, a decoding unit 201, an inverse quantization unit 202, an inverse transform unit 203, a first intra-frame prediction unit 204, a second intra-frame prediction unit 205, an inter-frame prediction unit 206, an adder 207, an accumulation unit 208, and an image output unit 220.
[0021] The symbol input unit 210 is configured to acquire the coded information coded by the image coding device.
[0022] The decoding unit 201 is configured to decode control information and quantization values from the coded information input from the symbol input unit 210. For example, the decoding unit 201 is configured to output control information and quantization values by performing variable length decoding on such coded information.
[0023] Here, the quantization values are sent to the inverse quantization unit 202, and the control information is sent to the intra prediction unit 204 for the first frame, the intra prediction unit 205 for the second frame, and the inter prediction unit 206. Note that such control information includes information necessary for controlling the intra prediction unit 204 for the first frame, the intra prediction unit 205 for the second frame, the inter prediction unit 206, etc., and may include header information such as a sequence parameter set, a picture parameter set, a picture header, and a slice header.
[0024] The inverse quantization unit 202 is configured to inverse quantize the quantization values sent from the decoding unit 201 to obtain transform coefficients. Such transform coefficients are sent to the inverse transform unit 203.
[0025] The inverse transform unit 203 is configured to inverse transform the transform coefficients sent from the inverse quantization unit 202 to obtain a prediction residual. Such a prediction residual is sent to the adder 207.
[0026] The intra prediction unit 204 for the first frame is configured to generate a first prediction pixel for adding to the prediction residual by the adder 207 based on the decoded pixel obtained via the adder 207 and the control information decoded by the decoding unit 201. Such a first prediction pixel is sent to the adder 207 and the intra prediction unit 205 for the second frame.
[0027] The inter prediction unit 206 is configured to generate a third prediction pixel for adding to the prediction residual by the adder 207 based on the decoded pixel obtained by referring to the accumulation unit 208 and the control information decoded by the decoding unit 201. Such a third prediction pixel is sent to the adder 207.
[0028] The accumulation unit 208 is configured to cumulatively accumulate the decoded pixels sent from the adder 207. Such decoded pixels receive a reference from the inter-frame prediction unit 206 via the accumulation unit 208.
[0029] The adder 207 is configured to add the prediction residual sent from the inverse transformation unit 203 and any one of the first to third predicted pixels sent from the first intra-frame prediction unit 204, the second intra-frame prediction unit 205, and the inter-frame prediction unit 206 to obtain a decoded pixel. Such a decoded pixel is sent to the image output unit 220, the accumulation unit 208, the first intra-frame prediction unit 204, and the second intra-frame prediction unit 205.
[0030] The image output unit 220 outputs the decoded pixel sent from the adder 207.
[0031] (Second Intra-Frame Prediction Unit 205) Hereinafter, an example of a method for deriving the second predicted pixel by the second intra-frame prediction unit 205 will be described.
[0032] The role of the second intra-frame prediction unit 205 is to derive a block vector (hereinafter referred to as BV) for the block to be decoded and predict the pixels of the block referred to by the BV in order to compensate the block to be decoded with high accuracy in the subsequent adder 207 (second intra-frame prediction).
[0033] Examples of the second intra-frame prediction include intra-template matching prediction (hereinafter referred to as IntraTMP) disclosed in Non-Patent Document 2.
[0034] When performing IntraTMP, the second in-frame prediction unit 205 uses decoded pixels adjacent to the block to be decoded within the same frame (the four lines adjacent to the left and above the block to be decoded shown in FIG. 2, the four lines adjacent to the left of the block to be decoded, or the four lines adjacent to the above of the block to be decoded) as templates, and searches for coordinates within the same frame where the templates match or are similar.
[0035] The second in-frame prediction unit 205 sets the displacement amount to such coordinates as BV, and refers to the block displaced from the block to be decoded and encoded by BV as the predicted pixel of the block to be decoded and encoded.
[0036] A plurality of BVs are constructed as a BV candidate list, and the BV is specified by signaling the index corresponding to the BV candidate in such BV candidate list.
[0037] When constructing the BV candidate list, the second in-frame prediction unit 205 calculates the template matching cost (sum of absolute differences (SAD), hereinafter referred to as cost) for each BV, and adds the BV with a small such cost to the BV candidate list.
[0038] In order to shorten the processing time for constructing the BV candidate list, the second in-frame prediction unit 205 compares the cost at that time with the threshold TH1 for each line of the template, and when such cost becomes equal to or greater than the threshold TH1 (first threshold), interrupts the calculation of such cost and does not add the BV candidate corresponding to such cost to the BV candidate list.
[0039] Conversely, when the second in-frame prediction unit 205 finishes calculating the costs of all lines of the template, it adds only the BV candidates with costs smaller than the threshold TH1 to the BV candidate list.
[0040] At this time, when the upper limit number of the BV candidate list is exceeded due to the addition of BV candidates, the second in-frame prediction unit 205 excludes the BV candidate with the maximum cost.
[0041] Such a threshold TH1 is set to a fixed value until the BV candidate list is satisfied, and after the BV candidate list is satisfied, it is set to the maximum cost in such a BV candidate list.
[0042] That is, each time a BV candidate with a small cost is added to the list, the threshold TH1 is updated and set to a smaller value.
[0043] According to the above configuration, when calculating the cost of a BV candidate, if the cost exceeds the threshold TH1 during the template matching, the processing time can be shortened by interrupting the template matching process at the point of exceeding.
[0044] From the perspective of shortening the processing time, it is desirable to set a small threshold at the earliest stage of the template matching process so that it can be interrupted early.
[0045] Therefore, the intra - frame prediction unit 205 of the second frame preferentially evaluates whether to add coordinates where there is a high possibility that the finally selected BV exists (that is, coordinates where there is a high possibility that a BV exists) to the BV candidate list as BV candidates.
[0046] For example, since the BV is likely to have the same value as the BV of a block close to the block to be decoded, as shown in FIG. 3, the intra - frame prediction unit 205 of the second frame preferentially evaluates whether to add the BV of a block close to the block to be decoded to the BV candidate list of the block to be decoded.
[0047] In the example of FIG. 3, the coordinates where there is a high possibility that a BV exists indicate the coordinates pointed to when the BV of an adjacent block is applied to the block to be decoded.
[0048] Note that as the blocks close to the block to be decoded, as shown in FIG. 4, adjacent blocks of the block to be decoded, blocks within a preset range, etc. can be used.
[0049] Also, since the same BV is likely to repeatedly occur even in blocks at a distant location, the in-frame prediction unit 205 in the second frame may preferentially evaluate whether to add the BV selected in the past to the BV candidate list of the block to be decoded, as shown in FIG. 5.
[0050] Alternatively, assuming that the same texture is repeated, coordinates that match or are similar are often directly to the right (y component of the BV is 0) and directly above (x component of the BV is 0) the block to be decoded. Therefore, as shown in FIG. 6, the in-frame prediction unit 205 in the second frame may preferentially evaluate whether to add the coordinates directly to the right or / and directly above the block to be decoded to the BV candidate list of the block to be decoded.
[0051] As shown in FIG. 7, the in-frame prediction unit 205 in the second frame may preferentially evaluate whether to add the coordinates directly to the right or / and directly above the coordinates preferentially evaluated in FIG. 5 to the BV candidate list.
[0052] In any configuration, since the maximum cost in the BV candidate list is likely to be small, an effect of shortening the processing time can be obtained.
[0053] Furthermore, when the in-frame prediction unit 205 in the second frame divides the evaluation range of the BV into a plurality of regions and evaluates for each region whether to add each coordinate in the region to the BV candidate list, it evaluates whether to add the coordinates closer to the block to be decoded within the region to the BV candidate list.
[0054] For example, as shown in FIG. 8, when a plurality of regions are set, in region 1, since the coordinates in the lower right are closest to the block to be decoded, the in-frame prediction unit 205 in the second frame evaluates from the coordinates in the lower right to the coordinates in the upper left.
[0055] On the other hand, in region 2, since the coordinates in the lower left are closest to the block to be decoded, the in-frame prediction unit 205 in the second frame evaluates from the coordinates in the lower left to the coordinates in the upper right.
[0056] To simplify the process, the in - second - frame prediction unit 205 may perform such evaluation line by line.
[0057] That is, the in - second - frame prediction unit 205 may first evaluate the horizontal line including the lower - right coordinate of region 1, and then repeatedly evaluate the horizontal lines one line above from right to left.
[0058] Alternatively, the in - second - frame prediction unit 205 may first evaluate the vertical line including the lower - right coordinate of region 1, and then repeatedly evaluate the vertical lines one line to the left from bottom to top.
[0059] When the coordinate closest to the block to be decoded is not the vertex of the region, such as in region 4 or region 5, the in - second - frame prediction unit 205 may substitute with the closest vertex, or may re - divide the region to generate vertices.
[0060] For example, the in - second - frame prediction unit 205 may re - divide region 4 vertically, evaluate from the lower - right coordinate to the upper - left coordinate in the upper part, and evaluate from the upper - right coordinate to the lower - left coordinate in the lower part.
[0061] Note that the in - second - frame prediction unit 205 may fix the evaluation order of regions in advance, or may change it adaptively.
[0062] When the in - second - frame prediction unit 205 fixes the evaluation order of regions in advance, it is desirable to evaluate in order from the regions close to the block to be decoded.
[0063] When the in - second - frame prediction unit 205 adaptively changes the evaluation order of regions, it is desirable to evaluate in order from the regions with a large number of BVs selected in the past.
[0064] If, as a result of the preferential evaluation as described above, the minimum cost in the BV candidate list is smaller than a separately determined threshold TH2 (second threshold), the in - second - frame prediction unit 205 can shorten the processing time by interrupting the construction process of the BV candidate list.
[0065] The second intra-frame prediction unit 205 may use a fixed value or adaptively change the threshold value TH2.
[0066] When the second intra-frame prediction unit 205 adaptively changes the threshold value TH2, for example, it is desirable to make the threshold value TH2 for interrupting the construction process of the BV candidate list proportional to the length of the BV.
[0067] Alternatively, as shown in FIG. 8, when the second intra-frame prediction unit 205 divides the evaluation range of the BV into a plurality of regions and evaluates the BV for each region, it can determine whether to interrupt the construction process of the BV candidate list for each region.
[0068] For example, when the second intra-frame prediction unit 205 preferentially evaluates as described above, if the minimum cost of the BV belonging to such a region in the BV candidate list constructed before the processing of such a region is greater than a separately determined threshold value TH3, the processing time can be shortened by interrupting the evaluation process of such a region.
[0069] The second intra-frame prediction unit 205 may use a fixed value or adaptively change the threshold value TH3.
[0070] When the second intra-frame prediction unit 205 adaptively changes the threshold value TH3, in any case, it is desirable to make the threshold value TH3 (the third threshold value) for interrupting the construction process of the BV candidate list proportional to the distance to the region.
[0071] Further, the second intra-frame prediction unit 205 may hierarchically evaluate the BV. Hereinafter, an example in which the second intra-frame prediction unit 205 hierarchically searches and evaluates the BV in two levels will be shown.
[0072] For example, in the first level, the second intra-frame prediction unit 205 evaluates the BV at a coarse spatial pixel interval, selects BV candidates, and then in the second level, evaluates the BV at a dense spatial pixel interval to determine the BV.
[0073] Since the BV candidates at the second layer are limited to only the periphery of the BV selected at the first layer, the effect of shortening the processing time can be obtained.
[0074] At this time, the intra-frame prediction unit 205 of the second frame may change the coarseness or fineness of the spatial pixel interval even within the same layer. For example, the intra-frame prediction unit 205 of the second frame changes the coarseness or fineness of such a spatial pixel interval according to the content.
[0075] Specifically, in screen content such as a document, there may be coordinates that exactly match, so if there is a deviation of even one pixel, the SAD changes greatly. Therefore, it is desirable that the intra-frame prediction unit 205 of the second frame be set to evaluate the BV densely in screen content.
[0076] On the other hand, in camera-captured content, since the distribution of the SAD is relatively flat, it is desirable that the intra-frame prediction unit 205 of the second frame be set to evaluate the BV sparsely in camera-captured content.
[0077] Also, the intra-frame prediction unit 205 of the second frame may change the coarseness or fineness of the spatial pixel interval according to the block size without using explicit control information.
[0078] Since coordinates that match or are similar are more easily evaluated as the block size is smaller, it is desirable that the intra-frame prediction unit 205 of the second frame evaluate the BV at a finer spatial pixel interval as the block size is smaller.
[0079] Conversely, since the improvement in coding efficiency is large when coordinates that match or are similar can be evaluated as the block size is larger, the intra-frame prediction unit 205 of the second frame may evaluate the BV at a finer spatial pixel interval as the block size is larger.
[0080] In any case, the intra-frame prediction unit 205 of the second frame may select the top N BV candidates and use them for the evaluation of the BV at the second layer.
[0081] Alternatively, the second intra prediction unit 205 may set the coarseness or fineness of the spatial pixel interval for each profile and level.
[0082] Hereinafter, the control information decoded by the decoder 201 when performing the second intra prediction will be described.
[0083] The code input to the decoder 201 may include a sequence parameter set (SPS) that aggregates control information at the sequence level.
[0084] Also, such a code may include a picture parameter set (PPS) or a picture header (PH) that aggregates control information at the picture level. Such a code may also include a slice header (SH) that aggregates control information at the slice level.
[0085] Referring to FIG. 9, a method for setting an evaluation method at the sequence level will be described.
[0086] As shown in FIG. 3, in step S101, the decoder 201 determines whether sps_itmp_enabled_flag is 1 in the sequence parameter set.
[0087] sps_itmp_enabled_flag is a syntax for controlling the presence or absence of IntraTMP. When sps_itmp_enabled_flag is 1, it indicates that IntraTMP is valid, and when sps_itmp_enabled_flag is 0, it indicates that IntraTMP is invalid.
[0088] If sps_itmp_enabled_flag is 1, this operation proceeds to step S102, and if sps_itmp_enabled_flag is 0, this operation ends.
[0089] In step S102, the decoder 201 decodes sps_itmp_mode.
[0090] The sps_itmp_mode is a syntax for controlling the IntraTMP method.
[0091] By using sps_itmp_mode, the evaluation method according to image characteristics can be changed in sequence units, so an effect of minimizing the processing time can be expected.
[0092] For example, for a sequence composed of CG, the pixel distribution is often composed of the same values, so the evaluation of the BV of IntraTMP can be set to be dense, and for a sequence composed of natural images, since the pixel distribution is diverse, the evaluation of the BV of IntraTMP can be set to be sparse, and the minimization of the processing time can be achieved.
[0093] When setting the correction method in picture units, the decoding unit 201 similarly decodes the pps_itmp_enabled_flag and pps_itmp_mode using the picture parameter set or the picture header.
[0094] By using pps_itmp_mode, the evaluation method according to image characteristics can be changed in picture units, so an effect of minimizing the processing time can be expected.
[0095] For example, for a picture composed of CG, the pixel distribution is often composed of the same values, so the evaluation of the BV of IntraTMP can be set to be dense, and for a picture composed of natural images, since the pixel distribution is diverse, the evaluation of the BV of IntraTMP can be set to be sparse, and the minimization of the processing time can be achieved.
[0096] When setting the correction method in slice units, the decoding unit 201 similarly decodes the sh_itmp_enabled_flag and sh_itmp_mode using the slice header.
[0097] By using sh_itmp_mode, the evaluation method according to image characteristics can be changed in slice units, so an effect of minimizing the processing time can be expected.
[0098] For example, for a slice area composed of CG, the pixel distribution is often composed of the same value, so the evaluation of the BV of IntraTMP can be set to be dense. For a slice area composed of natural images, since the pixel distribution is diverse, the evaluation of the BV of IntraTMP can be set to be sparse, and the minimization of the processing time can be achieved.
[0099] By setting only at the upper layer, an increase in the amount of code can be suppressed, and by setting at the lower layer as well and giving priority to the setting at the lower layer, adaptive control can be achieved.
[0100] Alternatively, when the above evaluation method is set in advance, the decoding of such an evaluation method itself can be omitted.
[0101] In the above example, the setting method of the IntraTMP method is described in units of sequence, picture, or slice. However, without setting these, the method can be directly selected in units of blocks described later. In this case, an increase in the above header information can be avoided.
[0102] Hereinafter, with reference to FIG. 10, a method for controlling the evaluation of the BV of IntraTMP in units of blocks will be described.
[0103] As shown in FIG. 10, in step S201, the decoding unit 201 determines whether any of sps_itmp_enabled_flag, pps_itmp_enabled_flag, or sh_itmp_enabled_flag is 1.
[0104] If none of them is 1, this operation ends. If any of them is 1, this operation proceeds to step S202.
[0105] In step S202, the decoding unit 201 decodes cu_itmp_mode, which is a control signal of IntraTMP.
[0106] In step S203, the decoding unit 201 decodes cu_itmp_index, which is a control signal representing an index.
[0107] According to the present embodiment, by adaptively constructing a BV list in the decoding of the BV of IntraTMP, the construction process of the BV candidate list can be terminated early, so that the processing time can be shortened.
[0108] The above-described image decoding apparatus 200 may be realized by a program that causes a computer to execute each function (each process).
Industrial Applicability
[0109] According to the present embodiment, for example, since an improvement in overall service quality can be realized in moving image communication, it is possible to contribute to Goal 9 of the Sustainable Development Goals (SDGs) led by the United Nations, "Build resilient infrastructure, promote sustainable industrialization, and foster innovation."
Explanation of Signs
[0110] 200... Image decoding apparatus 201... Decoding unit 202... Inverse quantization unit 203... Inverse transformation unit 204... First intra-frame prediction unit 205... Second intra-frame prediction unit 206... Inter-frame prediction unit 207... Adder 208... Accumulation unit 210... Symbol input unit 220... Image output unit
Claims
1. An image decoding apparatus, comprising: a decoding unit that decodes control information and quantization values; an inverse quantization unit that inverse quantizes the quantization values to obtain transform coefficients; an inverse transform unit that inverse transforms the transform coefficients to obtain prediction residuals; a first intra-frame prediction unit that generates a first predicted pixel based on decoded pixels and the control information; a second intra-frame prediction unit that preferentially evaluates coordinates where a block vector is likely to exist based on the decoded pixels and the control information, and generates a second predicted pixel based on the evaluation result; an accumulation unit that accumulates the decoded pixels; an inter-frame prediction unit that generates a third predicted pixel based on the accumulated decoded pixels and the control information; and an adder that adds the prediction residual and the first to third predicted pixels to obtain the decoded pixels.
2. The image decoding apparatus according to claim 1, wherein the second intra-frame prediction unit evaluates, as the evaluation, whether to add a block vector of a block close to a block to be decoded to a block vector candidate list of the block to be decoded.
3. The image decoding apparatus according to claim 2, wherein the block close to the block to be decoded is a block adjacent to the block to be decoded.
4. The image decoding apparatus according to claim 2, wherein the block close to the block to be decoded is a block within a preset range.
5. The image decoding apparatus according to claim 1, wherein the second intra-frame prediction unit evaluates, as the evaluation, whether to add a previously selected block vector to a block vector candidate list of the block to be decoded.
6. The image decoding apparatus according to claim 1, wherein the second intra-frame prediction unit evaluates, as the evaluation, whether to add at least one of directly to the left and directly above the block to be decoded to a block vector candidate list of the block to be decoded.
7. The image decoding apparatus according to claim 1, wherein the second intra-frame prediction unit evaluates, as the evaluation, whether to add at least one of directly to the left and directly above the coordinates to a block vector candidate list of the block to be decoded.
8. When the second in-frame predictor divides the range for performing the evaluation into a plurality of regions and performs the evaluation for each region, it evaluates from coordinates closer to the block to be decoded within the region. The image decoding apparatus according to claim 1, characterized in that.
9. When the second in-frame predictor divides the range for performing the evaluation into a plurality of regions and evaluates whether to add each coordinate within the region to the block vector candidate list for each region, it adaptively changes the order of the regions for performing the evaluation. The image decoding apparatus according to claim 1, characterized in that.
10. When the second in-frame predictor divides the range for performing the evaluation into a plurality of regions and evaluates whether to add each coordinate within the region to the block vector candidate list for each region, it evaluates in order from the region with a large number of previously selected block vectors. The image decoding apparatus according to claim 1, characterized in that.
11. When the minimum cost in the block vector candidate list of the second in-frame predictor is smaller than a second threshold value determined separately, the image decoding apparatus according to claim 1, characterized in that it interrupts the evaluation.
12. When the second in-frame predictor divides the range for performing the evaluation into a plurality of regions and performs the evaluation for each region, it determines whether to interrupt the construction process of the block vector candidate list for each region. The image decoding apparatus according to claim 1, characterized in that.
13. When the minimum cost of the block vector belonging to the region in the block vector candidate list constructed before the processing of the region by the second in-frame predictor is larger than a third threshold value determined separately, the evaluation for the region is interrupted. The image decoding apparatus according to claim 12, characterized in that.
14. When the second in-frame predictor performs the evaluation hierarchically, it changes the coarseness and fineness of the evaluation using a spatial pixel interval for each hierarchy. The image decoding apparatus according to claim 1, characterized in that.
15. The second in-frame predictor changes the coarseness and fineness of the spatial pixel interval according to the block size. The image decoding apparatus according to claim 14, characterized in that.
16. The second in-frame predictor performs the evaluation with a denser spatial pixel interval as the block size is smaller. The image decoding apparatus according to claim 15, characterized in that.
17. The image decoding apparatus according to claim 15, wherein the second intra-frame prediction unit performs the evaluation at a denser spatial pixel interval as the block size is larger.
18. An image decoding method, comprising: a step of decoding control information and quantization values; a step of inverse quantizing the quantization values to obtain transform coefficients; a method of inverse-transforming the transform coefficients to obtain prediction residuals; a step of generating a first predicted pixel based on decoded pixels and the control information; a step of preferentially evaluating coordinates where a block vector is likely to exist based on the decoded pixels and the control information, and generating a second predicted pixel based on the evaluation result; a step of accumulating the decoded pixels; a step of generating a third predicted pixel based on the accumulated decoded pixels and the control information; and a step of adding the prediction residuals and the first to third predicted pixels to obtain the decoded pixels.
19. A program for causing a computer to function as an image decoding apparatus, wherein the image decoding apparatus includes a decoding unit that decodes control information and quantization values; an inverse quantization unit that inverse quantizes the quantization values to obtain transform coefficients; an inverse transform unit that inverse-transforms the transform coefficients to obtain prediction residuals; a first intra-frame prediction unit that generates a first predicted pixel based on decoded pixels and the control information; a second intra-frame prediction unit that preferentially evaluates coordinates where a block vector is likely to exist based on the decoded pixels and the control information, and generates a second predicted pixel based on the evaluation result; an accumulation unit that accumulates the decoded pixels; an inter-frame prediction unit that generates a third predicted pixel based on the accumulated decoded pixels and the control information; and an adder that adds the prediction residuals and the first to third predicted pixels to obtain the decoded pixels.
Citation Information
Patent Citations
Motion detection method of image processing device
JP2642160B2