Image encoder, image encoding method, and program
By using multiple block quantization transformation coefficient groups and multi-scale quantization matrix in image encoding and decoding devices, combined with inverse quantization and inverse transformation processing, efficient zero output of orthogonal transformation coefficients is achieved, solving the problem of low zero output efficiency in the prior art, and improving encoding efficiency and image quality.
Patent Information
- Application Number
- JP2025034404
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2039-03-11
AI Technical Summary
When implementing the zero-out method in the prior art, it is difficult to efficiently set some orthogonal transformation coefficients to zero, affecting the encoding efficiency.
By using multiple block quantization transformation coefficient groups in image encoding and decoding devices, using multi-scale quantization matrix, combined with inverse quantization and inverse transformation processing, efficient zero output of orthogonal transformation coefficients is achieved.
The efficiency of the zero-out method is improved, the calculation amount is reduced, and the encoding efficiency and image quality are improved.
Smart Images

Figure 2025074308000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to image coding technology. [Background technology]
[0002] The High Efficiency Video Coding (HEVC) coding method (hereafter referred to as HEVC) is known as a coding method for compressing moving images. In order to improve coding efficiency, HEVC adopts a basic block larger than the conventional macroblock (16x16 pixels). This large basic block is called a Coding Tree Unit (CTU) and its maximum size is 64x64 pixels. The CTU is further divided into subblocks, which are the units for prediction and transformation.
[0003] In addition, in HEVC, a quantization matrix is used to weight coefficients after orthogonal transformation (hereinafter, referred to as orthogonal transformation coefficients) according to frequency components. By using a quantization matrix, data of high frequency components, degradation of which is less noticeable to human vision, is reduced more than data of low frequency components, thereby making it possible to improve compression efficiency while maintaining image quality. JP 2013-38758 A (Patent Document 1) discloses a technology for encoding information indicating such a quantization matrix.
[0004] Recently, activities have been started to internationally standardize a more efficient coding method as a successor to HEVC. Specifically, the standardization of the Versatile Video Coding (VVC) coding method (hereinafter referred to as VVC) is being promoted by the Joint Video Experts Team (JVET) established by ISO / IEC and ITU-T. In this standardization, in order to improve coding efficiency, a new method (hereinafter referred to as zero-out) is being considered that reduces the amount of code by forcibly setting the orthogonal transform coefficients of high frequency components to 0 when the block size during orthogonal transform is large. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Patent Publication No. 2013-38758 Summary of the Invention [Problem to be solved by the invention]
[0006] Therefore, an object of the present invention is to more efficiently execute a method for forcibly setting some orthogonal transform coefficients to 0. [Means for solving the problem]
[0007] To solve the above problems, the image decoding apparatus of the present invention has the following configuration. That is, in an image decoding apparatus capable of decoding an image from a bit stream using a plurality of blocks including a first block of P×Q pixels (P and Q are integers) and a second block of N×M pixels (N is an integer satisfying N<P, and M is an integer satisfying M<Q), decoding means for decoding from the bit stream data corresponding to a first quantization transform coefficient group corresponding to the first block and data corresponding to a second quantization transform coefficient group corresponding to the second block; inverse quantization means for deriving a first transform coefficient group representing frequency components from the first quantization transform coefficient group using a first quantization matrix having N×M elements, and deriving a second transform coefficient group representing frequency components from the second quantization transform coefficient group using a second quantization matrix having N×M elements; inverse transform means for deriving a first prediction error group corresponding to the first block by performing an inverse transform process on the first transform coefficient group, and deriving a second prediction error group corresponding to the second block by performing an inverse transform process on the second transform coefficient group; and a reproduction unit for reproducing at least image data based on the first prediction error group and prediction image data, wherein the prediction image data can be derived using a prediction method combining intra prediction and inter prediction. When the block to be decoded is the first block, the inverse transform means derives N×Q intermediate values by multiplying the first transform coefficient group, which is N×M transform coefficients, by a matrix of M×Q, and further derives the first prediction error group, which is P×Q prediction errors from the first transform coefficient group, by multiplying a matrix of P×N by the N×Q intermediate values. The first quantization matrix having N×M elements is a quantization matrix that includes some elements in a third quantization matrix having R×S (R is an integer satisfying R≦N, and S is an integer satisfying S≦M) elements and does not include other elements in the third quantization matrix. The second quantization matrix having N×M elements is a quantization matrix that includes all elements in a fourth quantization matrix having R×S elements. The third quantization matrix is different from the fourth quantization matrix.The first quantization matrix is a quantization matrix consisting of the part of elements in the third quantization matrix except for an element corresponding to a DC component, the second quantization matrix is a quantization matrix consisting of all of the elements in the fourth quantization matrix except for an element corresponding to a DC component, and the inverse quantization means further uses a quantization parameter to derive a set of transform coefficients.
[0008] To solve the above problems, the image encoding apparatus of the present invention has the following configuration. That is, in an image encoding apparatus capable of encoding an image using a plurality of blocks including a first block of P×Q pixels (P and Q are integers) and a second block of N×M pixels (N is an integer satisfying N<P, and M is an integer satisfying M<Q), a first conversion coefficient group is derived by performing a conversion process on a first prediction error group corresponding to the first block, and a second conversion coefficient group is derived by performing a conversion process on a second prediction error group corresponding to the second block. A conversion means, a quantization means for quantizing the first conversion coefficient group using a first quantization matrix having N×M elements to derive a first quantized conversion coefficient group, and quantizing the second conversion coefficient group using a second quantization matrix having N×M elements to derive a second quantized conversion coefficient group, and an encoding means for encoding data corresponding to the first quantized conversion coefficient group corresponding to the first block and data corresponding to the second quantized conversion coefficient group corresponding to the second block. At least the first prediction error group can be derived using a prediction method that combines intra prediction and inter prediction. When the block to be encoded is the first block, the conversion means derives P×M intermediate values by multiplying the first prediction error group, which is P×Q prediction errors, by a Q×M matrix, and further derives the first conversion coefficient group, which is N×M conversion coefficients, from the first prediction error group by multiplying an N×P matrix by the P×M intermediate values. The first quantization matrix having N×M elements is a quantization matrix that includes some elements in a third quantization matrix having R×S (R is an integer satisfying R≦N, and S is an integer satisfying S≦M) elements and does not include other elements in the third quantization matrix. The second quantization matrix having N×M elements is a quantization matrix that includes all elements in a fourth quantization matrix having R×S elements. The third quantization matrix is different from the fourth quantization matrix. The first quantization matrix, except for the element corresponding to the DC component,the second quantization matrix is a quantization matrix composed of all of the elements in the fourth quantization matrix except for an element corresponding to a DC component, and the quantization means further uses a quantization parameter to derive a set of quantized transform coefficients. Effect of the Invention
[0009] An object of the present invention is to more efficiently execute a method for forcibly setting some orthogonal transform coefficients to 0. [Brief description of the drawings]
[0010] [Figure 1] FIG. 1 is a block diagram showing a configuration of an image encoding device according to a first embodiment. [Diagram 2] FIG. 11 is a block diagram showing a configuration of an image decoding device according to a second embodiment. [Diagram 3] 4 is a flowchart showing an image encoding process in the image encoding device according to the first embodiment. [Figure 4] 11 is a flowchart showing an image decoding process in the image decoding device according to the second embodiment. [Diagram 5] FIG. 1 is a block diagram showing an example of the hardware configuration of a computer applicable to an image encoding device or an image decoding device of the present invention. [Figure 6] FIG. 2 is a diagram showing an example of a bit stream output in the first embodiment. [Figure 7] FIG. 2 is a diagram showing an example of sub-block division used in the first and second embodiments. [Figure 8] FIG. 4 is a diagram showing an example of a quantization matrix used in the first and second embodiments. [Figure 9] FIG. 4 is a diagram showing a method of scanning each element of a quantization matrix used in the first and second embodiments. [Figure 10] FIG. 11 is a diagram showing a difference value matrix of a quantization matrix generated in the first and second embodiments. [Figure 11] FIG. 13 is a diagram showing an example of a coding table used for coding a difference value of a quantization matrix. [Figure 12] FIG. 11 is a diagram showing another example of the quantization matrix used in the first and second embodiments. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0011] An embodiment of the present invention will be described with reference to the accompanying drawings. Note that the configurations shown in the following embodiments are merely examples, and the present invention is not limited to the configurations described in the following embodiments. Note that names such as basic block, sub-block, quantization matrix, and base quantization matrix are names used for convenience in each embodiment, and other names may be used as appropriate as long as their meanings do not change. For example, basic blocks and sub-blocks may be called basic units and sub-units, or simply blocks and units. In the following description, a rectangle is a quadrangle with four interior angles that are right angles and two diagonals that are equal in length, as is generally defined. A square is a quadrangle with all four corners and all four sides that are equal, as is generally defined. In other words, a square is a type of rectangle.
[0012] <Embodiment 1> Hereinafter, an embodiment of the present invention will be described with reference to the drawings.
[0013] First, zeroing out will be described in more detail. As described above, zeroing out is a process of forcibly setting some of the orthogonal transform coefficients of a block to be coded to 0. For example, a block of 64×64 pixels in an input image (picture) is assumed to be a block to be coded. In this case, the size of the orthogonal transform coefficients is also 64×64. Zeroing out is a process of coding some of the 64×64 orthogonal transform coefficients by regarding them as 0 even if they have a value other than 0 as a result of orthogonal transform. For example, low frequency components corresponding to a predetermined range in the upper left corner including DC components in two-dimensional orthogonal transform coefficients are not forcibly set to 0, and orthogonal transform coefficients corresponding to frequency components higher than these low frequency components are always set to 0.
[0014] Next, the image coding apparatus of this embodiment will be described. Fig. 1 is a block diagram showing the image coding apparatus of this embodiment. In Fig. 1, reference numeral 101 denotes a terminal for inputting image data.
[0015] Reference numeral 102 denotes a block dividing unit that divides an input image into a plurality of basic blocks and outputs an image in units of basic blocks to a subsequent stage.
[0016] Reference numeral 103 denotes a quantization matrix storage unit that generates and stores a quantization matrix. Here, the quantization matrix is used to weight the quantization process for the orthogonal transform coefficients according to the frequency components. In the quantization process described later, the quantization step for each orthogonal transform coefficient is weighted by, for example, multiplying a scale value (quantization scale) based on a reference parameter value (quantization parameter) by the value of each element in the quantization matrix.
[0017] There is no particular limitation on the method of generating the quantization matrix stored by the quantization matrix storage unit 110. For example, the user may input information indicating the quantization matrix, or the image encoding device may calculate the quantization matrix from the characteristics of the input image. Also, a previously specified initial value may be used. In this embodiment, in addition to the 8×8 base quantization matrix shown in FIG. 8(a), two types of 32×32 two-dimensional quantization matrices shown in FIG. 8(b) and (c) are generated and stored by enlarging the base quantization matrix. The quantization matrix in FIG. 8(b) is a 32×32 quantization matrix obtained by enlarging each element of the 8×8 base quantization matrix in FIG. 8(a) by four times in the vertical and horizontal directions. On the other hand, the quantization matrix in FIG. 8(c) is a 32×32 quantization matrix obtained by enlarging each element of the upper left 4×4 part of the base quantization matrix in FIG. 8(a) by eight times in the vertical and horizontal directions.
[0018] As described above, the base quantization matrix is a quantization matrix used not only for quantization in 8×8 pixel subblocks, but also for creating a quantization matrix of a size larger than the base quantization matrix. Note that the size of the base quantization matrix is assumed to be 8×8, but is not limited to this size. Also, a different base quantization matrix may be used depending on the size of the subblock. For example, when three types of subblocks, 8×8, 16×16, and 32×32, are used, three types of base quantization matrices corresponding to the respective types may be used.
[0019] A prediction unit 104 determines subblock division for image data in units of basic blocks. In other words, it determines whether or not to divide a basic block into subblocks, and if so, how to divide it. If not divided into subblocks, the subblocks will have the same size as the basic blocks. The subblocks may be square or may be rectangular other than a square.
[0020] Then, the prediction unit 104 performs intra-prediction, which is a prediction within a frame, or inter-prediction, which is a prediction between frames, on a sub-block basis to generate predicted image data.
[0021] For example, the prediction unit 104 selects a prediction method to be performed on one subblock from intra prediction and inter prediction, performs the selected prediction, and generates predicted image data for the subblock. However, the prediction method used is not limited to these, and a prediction that combines intra prediction and inter prediction may be used.
[0022] Furthermore, the prediction unit 104 calculates and outputs a prediction error from the input image data and the predicted image data. For example, the prediction unit 104 calculates the difference between each pixel value of a subblock and each pixel value of the predicted image data generated by predicting the subblock, and calculates it as a prediction error.
[0023] The prediction unit 104 also outputs information necessary for prediction, such as information indicating the division state of the sub-block, a prediction mode indicating a prediction method for the sub-block, information such as a motion vector, etc., together with the prediction error. Hereinafter, the information necessary for prediction is collectively referred to as prediction information.
[0024] Reference numeral 105 denotes a transform / quantization unit. The transform / quantization unit 105 performs orthogonal transform on the prediction error calculated by the prediction unit 104 in units of sub-blocks to obtain orthogonal transform coefficients representing each frequency component of the prediction error. The transform / quantization unit 105 then performs quantization using the quantization matrix stored in the quantization matrix holding unit 103 and the quantization parameter to obtain quantization coefficients that are quantized orthogonal transform coefficients. Note that the function of performing the orthogonal transform and the function of performing the quantization may be configured separately.
[0025] Reference numeral 106 denotes an inverse quantization and inverse transform unit. The inverse quantization and inverse transform unit 106 inverse quantizes the quantization coefficients output from the transform and quantization unit 105 using the quantization matrix stored in the quantization matrix storage unit 103 and the quantization parameter to reproduce the orthogonal transform coefficients. The inverse quantization and inverse transform unit 106 then performs inverse orthogonal transform to reproduce the prediction error. The process of reproducing (deriving) the orthogonal transform coefficients using the quantization matrix and the quantization parameter in this manner is referred to as inverse quantization. Note that the function of performing inverse quantization and the function of performing inverse quantization may be configured separately. Information for the image decoding device to derive the quantization parameter is also coded into a bit stream by the coding unit 110.
[0026] Reference numeral 108 denotes a frame memory for storing the reproduced image data.
[0027] An image reproduction unit 107 generates predicted image data based on the prediction information output from the prediction unit 104 by appropriately referring to a frame memory 108, and generates and outputs reproduced image data from this and the input prediction error.
[0028] An in-loop filter unit 109 performs in-loop filter processing such as a deblocking filter or a sample adaptive offset on a reconstructed image, and outputs a filtered image.
[0029] An encoding unit 110 encodes the quantization coefficients output from the transform / quantization unit 105 and the prediction information output from the prediction unit 104 to generate and output coded data.
[0030] A quantization matrix encoding unit 113 encodes the base quantization matrix output from the quantization matrix storage unit 103, and generates and outputs quantization matrix code data for use by the image decoding device to derive the base quantization matrix.
[0031] An integrated coding unit 111 generates header code data using the quantization matrix code data output from the quantization matrix coding unit 113. The integrated coding unit 111 further combines the header code data with the code data output from the coding unit 110 to form a bit stream and outputs it.
[0032] Reference numeral 112 denotes a terminal, which outputs the bit stream generated by the integrated encoding unit 111 to the outside.
[0033] The image coding operation in the image coding device will be described below. In this embodiment, moving image data is input in units of frames. Furthermore, for the sake of explanation, this embodiment will be described assuming that the block division unit 101 divides the image into basic blocks of 64×64 pixels, but is not limited to this. For example, a block of 128×128 pixels or a block of 32×32 pixels may be used as the basic block.
[0034] Prior to encoding an image, an image encoding device generates and encodes a quantization matrix. In the following description, as an example, the horizontal direction in the quantization matrix 800 and each block is defined as the x coordinate and the vertical direction is defined as the y coordinate, with the rightward direction being positive and the downward direction being positive, respectively. The coordinates of the upper left element in the quantization matrix 800 are defined as (0,0). That is, the coordinates of the lower right element of an 8×8 base quantization matrix are (7,7). The coordinates of the lower right element of a 32×32 quantization matrix are (31,31).
[0035] First, the quantization matrix storage unit 103 generates a quantization matrix. The quantization matrix is generated according to the size of the subblock, the size of the orthogonal transform coefficient to be quantized, and the type of prediction method. In this embodiment, first, an 8×8 base quantization matrix used for generating a quantization matrix described later and shown in FIG. 8(a) is generated. Next, this base quantization matrix is expanded to generate two types of 32×32 quantization matrices shown in FIG. 8(b) and FIG. 8(c). The quantization matrix in FIG. 8(b) is a 32×32 quantization matrix expanded four times by repeating each element of the 8×8 base quantization matrix in FIG. 8(a) four times in the vertical and horizontal directions.
[0036] That is, in the example shown in Fig. 8(b), each element in the 32x32 quantization matrix whose x coordinate is in the range of 0 to 3 and whose y coordinate is in the range of 0 to 3 is assigned the value of 1, which is the value of the upper left element in the base quantization matrix. Also, each element in the 32x32 quantization matrix whose x coordinate is in the range of 28 to 31 and whose y coordinate is in the range of 28 to 31 is assigned the value of 15, which is the value of the lower right element in the base quantization matrix. In the example of Fig. 8(b), all of the values of each element in the base quantization matrix are assigned to one of the elements of the 32x32 quantization matrix.
[0037] On the other hand, the quantization matrix in FIG. 8(c) is a 32×32 quantization matrix expanded by repeating each element in the upper left 4×4 part of the base quantization matrix in FIG. 8(a) eight times in the vertical and horizontal directions.
[0038] That is, in the example shown in Fig. 8(c), each element in the 32x32 quantization matrix whose x coordinate is in the range of 0 to 7 and whose y coordinate is in the range of 0 to 7 is assigned the value of 1, which is the value of the upper left element in the upper left 4x4 part of the base quantization matrix. Also, each element in the 32x32 quantization matrix whose x coordinate is in the range of 24 to 31 and whose y coordinate is in the range of 24 to 31 is assigned the value of 7, which is the value of the lower right element in the upper left 4x4 part of the base quantization matrix. In the example of Fig. 8(c), only the element values corresponding to the upper left 4x4 part (x coordinate range of 0 to 3 and y coordinate range of 0 to 3) of the values of each element in the base quantization matrix are assigned to each element of the 32x32 quantization matrix.
[0039] However, the quantization matrix to be generated is not limited to this, and when the size of the orthogonal transform coefficients to be quantized exists other than 32×32, a quantization matrix corresponding to the size of the orthogonal transform coefficients to be quantized, such as 16×16, 8×8, or 4×4, may be generated. The method of determining the base quantization matrix and each element constituting the quantization matrix is not particularly limited. For example, a predetermined initial value may be used, or each element may be set individually. Also, the quantization matrix may be generated according to the characteristics of the image.
[0040] The quantization matrix storage unit 103 stores the base quantization matrix and the quantization matrix thus generated. FIG. 8(b) shows an example of a quantization matrix used for quantizing an orthogonal transform coefficient corresponding to a 32×32 subblock described later, and FIG. 8(c) shows an example of a quantization matrix used for quantizing an orthogonal transform coefficient corresponding to a 64×64 subblock described later. The bold frame 800 represents a quantization matrix. For ease of explanation, each is assumed to be configured for 1024 pixels of 32×32, and each square in the bold frame represents each element constituting the quantization matrix. In this embodiment, the three types of quantization matrices shown in FIG. 8(b) and (c) are stored in a two-dimensional shape, but the elements in the quantization matrix are not limited to this. In addition, it is also possible to store multiple quantization matrices for the same prediction method depending on the size of the orthogonal transform coefficient to be quantized or whether the encoding target is a luminance block or a chrominance block. Generally, a quantization matrix realizes quantization processing that corresponds to the characteristics of human vision, so that the elements in the low-frequency part corresponding to the upper left part of the quantization matrix are small and the elements in the high-frequency part corresponding to the lower right part are large, as shown in Figures 8(b) and (c).
[0041] The quantization matrix encoding unit 113 sequentially reads out each element of the base quantization matrix stored in a two-dimensional form from the quantization matrix storage unit 106, scans each element to calculate a difference, and arranges each difference in a one-dimensional matrix. In this embodiment, the base quantization matrix shown in FIG. 8(a) uses the scanning method shown in FIG. 9, and calculates the difference between each element and the previous element in the scanning order. For example, the 8×8 base quantization matrix shown in FIG. 8(a) is scanned by the scanning method shown in FIG. 9, and after the first element 1 located in the upper left, the element 2 located immediately below it is scanned, and the difference +1 is calculated. In addition, the encoding of the first element of the quantization matrix (1 in this embodiment) is calculated by calculating the difference from a predetermined initial value (e.g., 8), but of course this is not limited to this, and the difference from an arbitrary value or the value of the first element itself may be used.
[0042] In this manner, in this embodiment, the base quantization matrix in Fig. 8(a) is scanned using the scanning method in Fig. 9, and the difference matrix shown in Fig. 10 is generated. The quantization matrix encoding unit 113 further encodes the difference matrix to generate quantization matrix code data. In this embodiment, the encoding is performed using the encoding table shown in Fig. 11(a), but the encoding table is not limited to this, and for example, the encoding table shown in Fig. 11(b) may be used. The quantization matrix code data generated in this manner is output to the integrated encoding unit 111 at the subsequent stage.
[0043] Returning to FIG. 1, the integrated coding unit 111 codes header information necessary for coding image data, and integrates the coded data of the quantization matrix.
[0044] Next, the image data is coded. One frame of image data is inputted from a terminal 101 to a block division unit .
[0045] The block division unit 102 divides the input image data into a plurality of basic blocks, and outputs an image in units of a basic block to a prediction unit 104. In this embodiment, it is assumed that an image in units of a basic block of 64×64 pixels is output.
[0046] The prediction unit 104 performs prediction processing on the image data in units of basic blocks input from the block division unit 102. Specifically, the prediction unit 104 determines sub-block division for further dividing the basic blocks into smaller sub-blocks, and further determines a prediction mode such as intra prediction or inter prediction for each sub-block.
[0047] FIG. 7 shows an example of a subblock division method. A basic block is shown in a bold frame 700. For ease of explanation, the basic block is assumed to be a 64×64 pixel structure, and each rectangle in the bold frame represents a subblock. FIG. 7(b) shows an example of a square subblock division of a quadtree, in which a basic block of 648×64 pixels is divided into subblocks of 32×32 pixels. Meanwhile, FIG. 7(c)-(f) show an example of a rectangular subblock division, in which the basic block is divided into rectangular subblocks of 32×64 pixels in length in FIG. 7(c) and 64×32 pixels in width in FIG. 7(d). Also, in FIG. 7(e) and (f), the basic block is divided into rectangular subblocks at a ratio of 1:2:1. In this way, not only squares but also rectangular subblocks other than squares are used for the encoding process. Also, the basic block may be further divided into a plurality of square blocks, and the subblock division may be performed based on the divided square blocks. In other words, the size of the basic block is not limited to 64×64 pixels, and basic blocks of a variety of sizes may be used.
[0048] In this embodiment, only quadtree division as shown in FIG. 7(a) and FIG. 7(b) is used, in which the basic block of 64×64 pixels is not divided, but the subblock division method is not limited to this. Tripartite division as shown in FIG. 7(e) and (f) or binary tree division as shown in FIG. 7(c) and FIG. 7(d) may be used. When subblock division other than that shown in FIG. 7(a) and FIG. 7(b) is used, a quantization matrix corresponding to the subblock to be used is generated in the quantization matrix storage unit 103. When a new base quantization matrix corresponding to the generated quantization matrix is also generated, the new base quantization matrix is also coded in the quantization matrix coding unit 113.
[0049] Further, the prediction method by the prediction unit 194 used in this embodiment will be described in more detail. In this embodiment, as an example, two types of prediction methods, intra prediction and inter prediction, are used. Intra prediction generates predicted pixels of the encoding target block using encoded pixels located spatially around the encoding target block, and also generates information on an intra prediction mode indicating the intra prediction method used among intra prediction methods such as horizontal prediction, vertical prediction, and DC prediction. Inter prediction generates predicted pixels of the encoding target block using encoded pixels of a frame that is temporally different from the encoding target block, and also generates motion information indicating a reference frame, a motion vector, and the like. As described above, the prediction unit 194 may use a prediction method that combines intra prediction and inter prediction.
[0050] Prediction image data is generated from the determined prediction mode and encoded pixels, and a prediction error is generated from the input image data and the prediction image data, and output to the transformation and quantization unit 105. Information such as sub-block division and prediction mode is output to the encoding unit 110 and image reproduction unit 107 as prediction information.
[0051] The transform / quantization unit 105 performs orthogonal transform / quantization on the input prediction error to generate quantized coefficients. First, an orthogonal transform process corresponding to the size of the subblock is performed to generate orthogonal transform coefficients, and then the orthogonal transform coefficients are quantized using a quantization matrix stored in the quantization matrix holding unit 103 according to the prediction mode to generate quantized coefficients. The orthogonal transform / quantization process will be described in more detail below.
[0052] When the 32×32 subblock division shown in FIG. 7(b) is selected, the 32×32 prediction error is subjected to an orthogonal transform using a 32×32 orthogonal transform matrix to generate 32×32 orthogonal transform coefficients. Specifically, a 32×32 orthogonal transform matrix, such as a discrete cosine transform (DCT), is multiplied by the 32×32 prediction error to calculate intermediate coefficients in a 32×32 matrix. The intermediate coefficients in a 32×32 matrix are further multiplied by the transpose of the above-mentioned 32×32 orthogonal transform matrix to generate 32×32 orthogonal transform coefficients. The 32×32 orthogonal transform coefficients thus generated are quantized using the 32×32 quantization matrix and quantization parameters shown in FIG. 8(b) to generate 32×32 quantized coefficients. Since there are four 32×32 subblocks in a 64×64 basic block, the above process is repeated four times.
[0053] On the other hand, when the 64×64 division state (no division) shown in Fig. 7(a) is selected, a 32×64 orthogonal transform matrix generated by thinning out odd-numbered rows (hereinafter referred to as odd rows) in the 64×64 orthogonal transform matrix is used for the 64×64 prediction error. In other words, 32×32 orthogonal transform coefficients are generated by performing an orthogonal transform using the 32×64 orthogonal transform matrix generated by thinning out the odd rows.
[0054] Specifically, first, odd-numbered rows are thinned out from a 64×64 orthogonal transform matrix to generate a 64×32 orthogonal transform matrix. Then, this 64×32 orthogonal transform matrix is multiplied by a 64×64 prediction error to generate intermediate coefficients in a 64×32 matrix. This 64×32 intermediate coefficient is multiplied by a 32×64 transposed matrix obtained by transposing the above-mentioned 64×32 orthogonal transform matrix to generate 32×32 orthogonal transform coefficients. Then, the transform / quantization unit 105 performs zeroing out by setting the generated 32×32 orthogonal transform coefficients to the coefficients in the upper left part (x coordinate range 0 to 31 and y coordinate range 0 to 31) of the 64×64 orthogonal transform coefficients and setting the rest to 0.
[0055] In this embodiment, the orthogonal transform is performed on the 64×64 prediction error using a 64×32 orthogonal transform matrix and a 32×64 transposed matrix obtained by transposing the 64×32 orthogonal transform matrix. The zero-out is performed by generating 32×32 orthogonal transform coefficients in this manner. This allows 32×32 orthogonal transform coefficients to be generated with less computational effort than a method in which some of the 64×64 orthogonal transform coefficients generated by performing a 64×64 orthogonal transform are forcibly set to 0 even if the value is not 0. In other words, the computational effort in the orthogonal transform can be reduced compared to a case in which an orthogonal transform is performed using a 64×64 orthogonal transform matrix, and the orthogonal transform coefficients to be zeroed out as a result are coded as 0 regardless of whether they are 0 or not. It should be noted that the computational effort can be reduced by using a method of calculating 32×32 orthogonal transform coefficients from 64×64 prediction errors using orthogonal transform coefficients, but the method of zeroing out is not limited to this method and various methods can be used.
[0056] Furthermore, when zeroing out is performed, information indicating that the orthogonal transform coefficients in the range subject to zeroing out are 0 may be coded, or information (such as a flag) indicating that zeroing out has been performed may be coded. By decoding such information, the image decoding device can decode each block by regarding the target of zeroing out as 0.
[0057] Next, the transform / quantization unit 105 quantizes the 32×32 orthogonal transform coefficients generated in this manner using the 32×32 quantization matrix shown in FIG. 8(c) and the quantization parameter to generate 32×32 quantized coefficients.
[0058] In this embodiment, the quantization matrix of FIG. 8(b) is used for 32×32 orthogonal transform coefficients corresponding to 32×32 sub-blocks, and the quantization matrix of FIG. 8(c) is used for 32×32 orthogonal transform coefficients corresponding to 64×64 sub-blocks. In other words, the quantization matrix of FIG. 8(b) is used for 32×32 orthogonal transform coefficients that have not been zeroed out, and the quantization matrix of FIG. 8(c) is used for 32×32 orthogonal transform coefficients corresponding to 64×64 sub-blocks that have been zeroed out. However, the quantization matrix used is not limited to this. The generated quantization coefficients are output to the encoding unit 110 and the inverse quantization and inverse transform unit 106.
[0059] The inverse quantization and inverse transform unit 106 inverse quantizes the input quantized coefficients using the quantization matrix stored in the quantization matrix holding unit 103 and the quantization parameters to reproduce the orthogonal transform coefficients. The inverse quantization and inverse transform unit 106 then performs inverse orthogonal transform on the reproduced orthogonal transform coefficients to reproduce the prediction error. As with the transform and quantization unit 105, the inverse quantization process uses a quantization matrix corresponding to the size of the sub-block to be coded. The inverse quantization and inverse orthogonal transform process performed by the inverse quantization and inverse transform unit 106 will be described more specifically below.
[0060] When the 32×32 subblock division in FIG. 7(b) is selected, the inverse quantization / inverse transform unit 106 inverse quantizes the 32×32 quantized coefficients generated by the transform / quantization unit 105 using the quantization matrix in FIG. 8(b) to reproduce 32×32 orthogonal transform coefficients. The inverse quantization / inverse transform unit 106 then multiplies the above-mentioned 32×32 transposed matrix by the 32×32 orthogonal transform to calculate intermediate coefficients in a 32×32 matrix. The inverse quantization / inverse transform unit 106 then multiplies the intermediate coefficients in the 32×32 matrix by the above-mentioned 32×32 orthogonal transform matrix to reproduce 32×32 prediction errors. The same process is performed for each 32×32 subblock. On the other hand, when no division is selected as shown in FIG. 7(a), the 32×32 quantization coefficients generated by the transform / quantization unit 105 are inversely quantized using the quantization matrix of FIG. 8(c) to reproduce 32×32 orthogonal transform coefficients. Then, the above-mentioned 32×64 transposed matrix is multiplied with the 32×32 orthogonal transform to calculate intermediate coefficients in a 32×64 matrix. The intermediate coefficients in the 32×64 matrix are multiplied with the above-mentioned 64×32 orthogonal transform matrix to reproduce 64×64 prediction errors. In this embodiment, the same quantization matrix as that used by the transform / quantization unit 105 is used according to the size of the subblock to perform the inverse quantization process. The reproduced prediction errors are output to the image reproduction unit 107.
[0061] The image reproduction unit 107 reproduces a predicted image by appropriately referring to data necessary for reproducing the predicted image stored in the frame memory 108 based on the prediction information input from the prediction unit 104. Then, image data is reproduced from the reproduced predicted image and the reproduced prediction error input from the inverse quantization and inverse transform unit 106, and is input to the frame memory 108 for storage.
[0062] The in-loop filter unit 109 reads the reconstructed image from the frame memory 108 and performs in-loop filtering such as deblocking filtering, etc. Then, the filtered image is input again to the frame memory 108 and stored again.
[0063] The coding unit 110 performs entropy coding on a block-by-block basis on the quantization coefficients generated by the transform / quantization unit 105 and the prediction information input from the prediction unit 104, to generate coded data. There is no particular specification as to the entropy coding method, but Golomb coding, arithmetic coding, Huffman coding, etc. can be used. The generated coded data is output to the integrated coding unit 111.
[0064] The integrated encoding unit 111 multiplexes the encoded data of the header and the encoded data input from the encoding unit 110 to form a bit stream. Finally, the bit stream is output from a terminal 112 to the outside.
[0065] FIG. 6(a) is an example of a bit stream output in the first embodiment. The sequence header includes the coded data of the base quantization matrix, and is composed of the coding results of each element. However, the position where the coded data of the base quantization matrix is coded is not limited to this, and it may be coded in the picture header section or other header section. In addition, when changing the quantization matrix in one sequence, it is also possible to update it by newly coding the base quantization matrix. In this case, all the quantization matrices may be rewritten, or it is also possible to change a part of it by specifying the size of the sub-block of the quantization matrix corresponding to the quantization matrix to be rewritten.
[0066] FIG. 3 is a flowchart showing the encoding process in the image encoding device according to the first embodiment.
[0067] First, prior to encoding an image, in step S301, the quantization matrix storage unit 103 generates and stores a two-dimensional quantization matrix. In this embodiment, the base quantization matrix shown in Fig. 8(a) and the quantization matrices shown in Fig. 8(b) and (c) generated from the base quantization matrix are generated and stored.
[0068] In step S302, the quantization matrix encoding unit 113 scans the base quantization matrix used to generate the quantization matrix in step S301, calculates the difference between each element that comes before and after in the scanning order, and generates a one-dimensional difference matrix. In this embodiment, the base quantization matrix shown in Fig. 8(a) is scanned using the scanning method in Fig. 9, and the difference matrix shown in Fig. 10 is generated. The quantization matrix encoding unit 113 further encodes the generated difference matrix to generate quantization matrix code data.
[0069] In step S303, the integrated encoding unit 111 encodes and outputs header information necessary for encoding image data together with the generated quantization matrix code data.
[0070] In step S304, the block division unit 102 divides the input image in units of frames into basic blocks of 64×64 pixels.
[0071] In step S305, the prediction unit 104 performs prediction processing on the basic block unit image data generated in step S304 using the prediction method described above, and generates prediction information such as sub-block division information and a prediction mode, and predicted image data. In this embodiment, two types of sub-block sizes are used: the 32×32 pixel sub-block division shown in Fig. 7(b) and the 64×64 pixel sub-block shown in Fig. 7(a). Furthermore, a prediction error is calculated from the input image data and the predicted image data.
[0072] In step S306, the transform / quantization unit 105 performs an orthogonal transform on the prediction error calculated in step S305 to generate orthogonal transform coefficients. The transform / quantization unit 105 then performs quantization using the quantization matrix and quantization parameter generated and held in step S301 to generate quantization coefficients. Specifically, the prediction error of the 32×32 pixel subblock in FIG. 7(b) is multiplied with a 32×32 orthogonal transform matrix and its transpose matrix to generate 32×32 orthogonal transform coefficients. Meanwhile, the prediction error of the 64×64 pixel subblock in FIG. 7(a) is multiplied with a 64×32 orthogonal transform matrix and its transpose matrix to generate 32×32 orthogonal transform coefficients. In this embodiment, the quantization matrix of Figure 8(b) is used for the orthogonal transform coefficients of the 32x32 sub-block in Figure 7(b), and the quantization matrix of Figure 8(c) is used for the orthogonal transform coefficients corresponding to the 64x64 sub-block in Figure 7(a), so that the 32x32 orthogonal transform coefficients are quantized.
[0073] In step S307, the inverse quantization and inverse transform unit 106 inverse quantizes the quantized coefficients generated in step S306 using the quantization matrix and the quantization parameters generated and held in step S301, to reproduce the orthogonal transform coefficients. Furthermore, the orthogonal transform coefficients are inversely transformed to reproduce the prediction error. In this step, the same quantization matrix as that used in step S306 is used, and the inverse quantization process is performed. Specifically, for the 32×32 quantized coefficients corresponding to the 32×32 pixel subblocks in FIG. 7(b), the inverse quantization process is performed using the quantization matrix in FIG. 8(b), to reproduce the 32×32 orthogonal transform coefficients. Then, the 32×32 orthogonal transform coefficients are multiplied by the 32×32 orthogonal transform matrix and its transpose matrix to reproduce the 32×32 pixel prediction error. On the other hand, the 32x32 quantized coefficients corresponding to the 64x64 pixel subblock in Fig. 7(a) are inverse quantized using the quantization matrix in Fig. 8(c) to reconstruct 32x32 orthogonal transform coefficients. These 32x32 orthogonal transform coefficients are then multiplied by a 64x32 orthogonal transform matrix and its transpose matrix to reconstruct 64x64 pixel prediction errors.
[0074] In step S308, the image reconstructor 107 reconstructs a predicted image based on the prediction information generated in step S305, and further reconstructs image data from the reconstructed predicted image and the prediction error generated in step S307.
[0075] In step S309, the encoding unit 110 encodes the prediction information generated in step S305 and the quantized coefficients generated in step S306 to generate coded data, and also generates a bit stream including other coded data.
[0076] In step S310, the image encoding device determines whether or not encoding of all basic blocks in the frame has been completed. If so, the image encoding device proceeds to step S311; if not, the image encoding device returns to step S304 for the next basic block.
[0077] In step S311, the in-loop filter unit 109 performs in-loop filter processing on the image data reproduced in step S308 to generate a filtered image, and then the process ends.
[0078] The above configuration and operation can reduce the amount of calculation while controlling quantization for each frequency component, thereby improving the subjective image quality. In particular, in step S305, the number of orthogonal transform coefficients is reduced, and quantization processing is performed using a quantization matrix corresponding to the reduced orthogonal transform coefficients, thereby controlling quantization for each frequency component while reducing the amount of calculation, thereby improving the subjective image quality. Furthermore, when the number of orthogonal transform coefficients is reduced and only the low frequency part is quantized and coded, an optimal quantization control for the low frequency part can be realized by using a quantization matrix that expands only the low frequency part of the base quantization matrix as shown in FIG. 8(c). Note that the low frequency part here is the range of x coordinate 0 to 3 and y coordinate 0 to 3 in the example of FIG. 8(c).
[0079] In this embodiment, in order to reduce the amount of code, only the base quantization matrix of FIG. 8(a) which is commonly used to generate the quantization matrices of FIG. 8(b) and FIG. 8(c) is coded. However, the quantization matrices of FIG. 8(b) and FIG. 8(c) themselves may be coded. In this case, since a unique value can be set for each frequency component of each quantization matrix, more detailed quantization control can be realized for each frequency component. It is also possible to set individual base quantization matrices for each of FIG. 8(b) and FIG. 8(c) and code each of the base quantization matrices. In this case, different quantization controls can be performed for the 32×32 orthogonal transform coefficients and the 64×64 orthogonal transform coefficients, and more detailed subjective image quality control can be realized. Furthermore, in this case, the quantization matrix corresponding to the 64×64 orthogonal transform coefficients may be enlarged by 4 times, instead of enlarging the upper left 4×4 part of the 8×8 base quantization matrix by 8 times. In this way, finer quantization control can be achieved even for 64×64 orthogonal transform coefficients.
[0080] Furthermore, in this embodiment, the quantization matrix for the 64×64 subblock using zero-out is uniquely determined, but it may be selectable by introducing an identifier. For example, FIG. 6(b) shows a case where a quantization matrix coding method information code is newly introduced to selectively code the quantization matrix for the 64×64 subblock using zero-out. For example, when the quantization matrix coding method information code indicates 0, an independent quantization matrix shown in FIG. 8(c) is used for the orthogonal transform coefficients corresponding to the 64×64 pixel subblock using zero-out. When the coding method information code indicates 1, the quantization matrix shown in FIG. 8(b) for the normal subblock not zeroed out is used for the 64×64 pixel subblock using zero-out. On the other hand, when the coding method information code indicates 2, all elements of the quantization matrix used for the 64×64 pixel subblock using zero-out are coded instead of the 8×8 base quantization matrix. This makes it possible to selectively realize a reduction in the amount of quantization matrix code and unique quantization control for the subblock using zero-out.
[0081] In addition, in this embodiment, the subblocks processed using zero-out are only 64×64, but the subblocks processed using zero-out are not limited to this. For example, among the orthogonal transform coefficients corresponding to the 32×64 or 64×32 subblocks shown in FIG. 7(c) or FIG. 7(b), the 32×32 orthogonal transform coefficients in the lower half or right half may be forcibly set to 0. In this case, only the 32×32 orthogonal transform coefficients in the upper half or left half are subject to quantization and encoding, and the quantization process is performed on the 32×32 orthogonal transform coefficients in the upper half or left half using a quantization matrix different from that in FIG. 8(b).
[0082] Furthermore, the value of the quantization matrix corresponding to the DC coefficient located at the upper left corner among the generated orthogonal transform coefficients, which is considered to have the greatest effect on image quality, may be set and coded separately from the values of the elements of the 8×8 base matrix. Figures 12(b) and 12(c) show an example in which the value of the element located at the upper left corner, which corresponds to the DC component, is changed compared to Figures 8(b) and 8(c). In this case, the quantization matrix shown in Figures 12(b) and 12(c) can be set by separately coding information indicating "2" located in the DC part in addition to the information of the base quantization matrix in Figure 8(a). This allows more fine quantization control to be performed on the DC component of the orthogonal transform coefficient, which has the greatest effect on image quality.
[0083] <Embodiment 2> 2 is a block diagram showing the configuration of an image decoding device according to a second embodiment of the present invention. In this embodiment, an image decoding device that decodes the encoded data generated in the first embodiment will be described as an example.
[0084] Reference numeral 201 denotes a terminal to which an encoded bit stream is input.
[0085] 202 is a separation decoding unit that separates the bit stream into information related to the decoding process and coded data related to coefficients, and decodes the coded data present in the header part of the bit stream. In this embodiment, the separation decoding unit 202 separates the quantization matrix code and outputs it to the subsequent stage. The separation decoding unit 202 performs the reverse operation of the integrated coding unit 111 in FIG. 1.
[0086] Reference numeral 209 denotes a quantization matrix decoding unit which decodes the quantization matrix code from the bit stream to reproduce the base quantization matrix, and further executes a process of generating each quantization matrix from the base quantization matrix.
[0087] Reference numeral 203 denotes a decoding unit, which decodes the coded data output from the separate decoding unit 202 and reproduces (derives) the quantization coefficients and prediction information.
[0088] 204 is an inverse quantization and inverse transform unit, which, like the inverse quantization and inverse transform unit 106 in Fig. 1, performs inverse quantization on the quantized coefficients using the reproduced quantization matrix and quantization parameter to obtain orthogonal transform coefficients, and further performs inverse orthogonal transform to reproduce prediction errors. Note that information for deriving the quantization parameter is also decoded from the bitstream by the decoding unit 203. Also, the function of performing inverse quantization and the function of performing inverse quantization may be configured separately.
[0089] Reference numeral 206 denotes a frame memory, which stores image data of reproduced pictures.
[0090] An image reproduction unit 205 generates predicted image data based on the input prediction information by appropriately referring to a frame memory 206. Then, the image reproduction unit 205 generates and outputs reproduced image data from the predicted image data and the prediction error reproduced by the inverse quantization and inverse transform unit 204.
[0091] An in-loop filter unit 207 performs in-loop filtering such as a deblocking filter on a reconstructed image, as in the in-loop filter unit 109 in Fig. 1, and outputs a filtered image.
[0092] Reference numeral 208 denotes a terminal for outputting the reproduced image data to the outside.
[0093] The image decoding operation in the image decoding device will be described below. In this embodiment, the bit stream generated in the first embodiment is input in units of frames (units of pictures).
[0094] In Fig. 2, a bit stream for one frame input from a terminal 201 is input to a separate decoding unit 202. The separate decoding unit 202 separates the bit stream into information related to the decoding process and coded data related to coefficients, and decodes the coded data present in the header of the bit stream. More specifically, the quantization matrix coded data is reproduced. In this embodiment, first, the quantization matrix coded data is extracted from the sequence header of the bit stream shown in Fig. 6(a) and output to the quantization matrix decoding unit 209. In this embodiment, the quantization matrix coded data corresponding to the base quantization matrix shown in Fig. 8(a) is extracted and output. Next, the coded data of the picture data in units of basic blocks is reproduced and output to the decoding unit 203.
[0095] The quantization matrix decoding unit 209 first decodes the input quantization matrix code data and reproduces the one-dimensional difference matrix shown in FIG. 10. In this embodiment, as in the first embodiment, the decoding is performed using the coding table shown in FIG. 11(a). However, the coding table is not limited to this, and other coding tables may be used as long as they are the same as those in the first embodiment. Furthermore, the quantization matrix decoding unit 209 reproduces the two-dimensional quantization matrix from the reproduced one-dimensional difference matrix. Here, the operation is the reverse of the operation of the quantization matrix encoding unit 113 in the first embodiment. That is, in this embodiment, the difference matrix shown in FIG. 10 reproduces and holds the base quantization matrix shown in FIG. 8(a) using the scanning method shown in FIG. 9. Specifically, the quantization matrix decoding unit 209 reproduces each element in the quantization matrix by sequentially adding each difference value in the difference matrix from the above-mentioned initial value. Then, the quantization matrix decoding unit 209 reproduces the two-dimensional quantization matrix by sequentially associating each of the reproduced one-dimensional elements with each of the elements of the two-dimensional quantization matrix according to the scanning method shown in FIG.
[0096] Furthermore, the quantization matrix decoding unit 209 enlarges this reconstructed base quantization matrix in the same manner as in the first embodiment, generating two types of 32×32 quantization matrices shown in Fig. 8(b) and Fig. 8(c). The quantization matrix in Fig. 8(b) is a 32×32 quantization matrix obtained by enlarging each element of the 8×8 base quantization matrix in Fig. 8(a) by four times in the vertical and horizontal directions.
[0097] On the other hand, the quantization matrix in Fig. 8(c) is a 32x32 quantization matrix enlarged by repeating each element of the 4x4 part in the upper left of the base quantization matrix in Fig. 8(a) eight times in the vertical and horizontal directions. However, the generated quantization matrix is not limited to this, and if the size of the quantization coefficients to be inverse quantized in the subsequent stage exists other than 32x32, a quantization matrix corresponding to the size of the quantization coefficients to be inverse quantized, such as 16x16, 8x8, or 4x4, may be generated. These generated quantization matrices are retained and used in the inverse quantization process in the subsequent stage.
[0098] The decoding unit 203 decodes the coded data from the bit stream and reproduces the quantization coefficients and prediction information. The size of the sub-block to be decoded is determined based on the decoded prediction information, and the reproduced quantization coefficients are output to the inverse quantization and inverse transform unit 204, and the reproduced prediction information is output to the image reproduction unit 205. In this embodiment, regardless of the size of the sub-block to be decoded, that is, whether it is 64×64 in FIG. 7(a) or 32×32 in FIG. 7(b), a 32×32 quantization coefficient is reproduced for each sub-block.
[0099] The inverse quantization and inverse transform unit 204 performs inverse quantization on the input quantized coefficients using the quantization matrix reproduced by the quantization matrix decoding unit 209 and the quantization parameters to generate orthogonal transform coefficients, and further performs inverse orthogonal transform to reproduce prediction errors. A more specific description of the inverse quantization and inverse orthogonal transform process is given below.
[0100] When the 32×32 subblock division in FIG. 7(b) is selected, the 32×3 quantization coefficients reproduced by the decoding unit 203 are inverse quantized using the quantization matrix in FIG. 8(b) to reproduce 32×32 orthogonal transform coefficients. Then, the above-mentioned 32×32 transposed matrix is multiplied with the 32×32 orthogonal transform to calculate intermediate coefficients in a 32×32 matrix. The intermediate coefficients in the 32×32 matrix are multiplied with the above-mentioned 32×32 orthogonal transform matrix to reproduce 32×32 prediction errors. The same process is performed for each 32×32 subblock.
[0101] On the other hand, when no division is selected as in Fig. 7(a), the 32x32 quantized coefficients reproduced by the decoding unit 203 are inverse quantized using the quantization matrix in Fig. 8(c) to reproduce 32x32 orthogonal transform coefficients. Then, the above-mentioned 32x64 transposed matrix is multiplied with the 32x32 orthogonal transform to calculate intermediate coefficients in a 32x64 matrix. This 32x64 intermediate coefficient matrix is multiplied with the above-mentioned 64x32 orthogonal transform matrix to reproduce 64x64 prediction errors.
[0102] The reconstructed prediction error is output to the image reconstruction unit 205. In this embodiment, the quantization matrix used in the inverse quantization process is determined according to the size of the subblock to be decoded, which is determined by the prediction information reconstructed by the decoding unit 203. That is, the quantization matrix in FIG. 8(b) is used in the inverse quantization process for each 32×32 subblock in FIG. 7(b), and the quantization matrix in FIG. 8(c) is used for each 64×64 subblock in FIG. 7(a). However, the quantization matrix used is not limited to this, and may be the same as the quantization matrix used in the transform / quantization unit 105 and the inverse quantization / inverse transform unit 106 in the first embodiment.
[0103] The image reproduction unit 205 appropriately refers to the frame memory 206 based on the prediction information input from the decoding unit 203, acquires data necessary for reproducing the predicted image, and reproduces the predicted image. In this embodiment, like the prediction unit 104 in the first embodiment, two types of prediction methods, intra prediction and inter prediction, are used. Also, as described above, a prediction method combining intra prediction and inter prediction may be used. Also, like the first embodiment, the prediction process is performed on a sub-block basis.
[0104] The specific prediction process is similar to that of the prediction unit 104 in the first embodiment, and therefore a description thereof will be omitted. The image reproduction unit 205 reproduces image data from a predicted image generated by the prediction process and a prediction error input from the inverse quantization and inverse transform unit 204. Specifically, the image reproduction unit 205 reproduces image data by adding the predicted image and the prediction error. The reproduced image data is stored in the frame memory 206 as appropriate. The stored image data is referred to as appropriate when predicting other sub-blocks.
[0105] 1, the in-loop filter unit 207 reads out the reconstructed image from the frame memory 206 and performs in-loop filtering such as deblocking filtering. The filtered image is then input back to the frame memory 206.
[0106] The reproduced image stored in the frame memory 206 is finally output to the outside from a terminal 208. The reproduced image is output to, for example, an external display device.
[0107] FIG. 4 is a flowchart showing an image decoding process in the image decoding device according to the second embodiment.
[0108] First, in step S401, the demultiplexing / decoding unit 202 demultiplexes the bit stream into information related to the decoding process and coded data related to coefficients, and decodes the coded data of the header portion. More specifically, it reproduces the quantization matrix coded data.
[0109] In step S402, the quantization matrix decoding unit 209 first decodes the quantization matrix code data reproduced in step S401 to reproduce the one-dimensional difference matrix shown in Fig. 10. Next, the quantization matrix decoding unit 209 reproduces a two-dimensional base quantization matrix from the reproduced one-dimensional difference matrix. Furthermore, the quantization matrix decoding unit 209 enlarges the reproduced two-dimensional base quantization matrix to generate a quantization matrix.
[0110] That is, in this embodiment, the quantization matrix decoding unit 209 reproduces the base quantization matrix shown in Fig. 8(a) by using the difference matrix shown in Fig. 10 and the scanning method shown in Fig. 9. Furthermore, the quantization matrix decoding unit 209 enlarges the reproduced base quantization matrix to generate and hold the quantization matrices shown in Fig. 8(b) and Fig. 8(c).
[0111] In step S403, the decoding unit 203 decodes the coded data separated in step S401 to reproduce the quantization coefficients and prediction information. Furthermore, the size of the subblock to be decoded is determined based on the decoded prediction information. In this embodiment, regardless of the size of the subblock to be decoded, i.e., 64×64 in FIG. 7(a) or 32×32 in FIG. 7(b), a 32×32 quantization coefficient is reproduced for each subblock.
[0112] In step S404, the inverse quantization and inverse transform unit 204 performs inverse quantization on the quantized coefficients using the quantization matrix reproduced in step S402 to obtain orthogonal transform coefficients, and further performs inverse orthogonal transform to reproduce prediction errors. In this embodiment, the quantization matrix used in the inverse quantization process is determined according to the size of the subblock to be decoded determined by the prediction information reproduced in step S403. That is, the quantization matrix in FIG. 8(b) is used in the inverse quantization process for each 32×32 subblock in FIG. 7(b), and the quantization matrix in FIG. 8(c) is used for each 64×64 subblock in FIG. 7(a). However, the quantization matrix used is not limited to this, and may be the same as the quantization matrix used in steps S306 and S307 in the first embodiment.
[0113] In step S405, the image reproducing unit 205 reproduces a predicted image from the prediction information generated in step S403. In this embodiment, two types of prediction methods, intra prediction and inter prediction, are used, as in step S305 in embodiment 1. Furthermore, image data is reproduced from the reproduced predicted image and the prediction error generated in step S404.
[0114] In step S406, the image decoding apparatus determines whether or not the decoding of all basic blocks in the frame has been completed. If the decoding has been completed, the image decoding apparatus proceeds to step S407, and if not, the image decoding apparatus returns to step S403 for the next basic block.
[0115] In step S407, the in-loop filter unit 207 performs in-loop filter processing on the image data reproduced in step S405 to generate a filtered image, and then the process ends.
[0116] With the above configuration and operation, it is possible to decode a bitstream in which quantization is controlled for each frequency component using a quantization matrix and subjective image quality is improved, even for subblocks in which only low-frequency orthogonal transform coefficients are quantized and coded, as generated in embodiment 1. In addition, for subblocks in which only low-frequency orthogonal transform coefficients are quantized and coded, a quantization matrix in which only the low-frequency part of the base quantization matrix as shown in Fig. 8(c) is enlarged is used, and a bitstream in which optimal quantization control is applied to the low-frequency part can be decoded.
[0117] In this embodiment, in order to reduce the amount of code, only the base quantization matrix in Fig. 8(a) that is commonly used to generate the quantization matrices in Fig. 8(b) and (c) is decoded, but the quantization matrices in Fig. 8(b) and (c) themselves may be decoded. In that case, a unique value can be set for each frequency component of each quantization matrix, so that a bitstream that realizes finer quantization control for each frequency component can be decoded.
[0118] It is also possible to configure a configuration in which separate base quantization matrices are set for each of Figures 8(b) and 8(c) and each base quantization matrix is coded. In this case, different quantization controls are performed for the 32x32 orthogonal transform coefficients and the 64x64 orthogonal transform coefficients, and a bitstream that realizes more precise control of subjective image quality can be decoded. Furthermore, in this case, the quantization matrix corresponding to the 64x64 orthogonal transform coefficients may be enlarged by a factor of four for the entire 8x8 base quantization matrix, instead of enlarging the upper left 4x4 portion of the 8x8 base quantization matrix by a factor of eight. In this way, more precise quantization control can be realized for the 64x64 orthogonal transform coefficients.
[0119] Furthermore, in this embodiment, the quantization matrix for the 64×64 sub-block using zero-out is uniquely determined, but it may be selectable by introducing an identifier. For example, FIG. 6(b) shows a case where the quantization matrix coding method information code is newly introduced to selectively code the quantization matrix for the 64×64 sub-block using zero-out. For example, when the quantization matrix coding method information code indicates 0, the independent quantization matrix shown in FIG. 8(c) is used for the quantization coefficients corresponding to the 64×64 sub-block using zero-out. When the coding method information code indicates 1, the quantization matrix shown in FIG. 8(b) for the normal sub-block not zeroed out is used for the 64×64 sub-block using zero-out. On the other hand, when the coding method information code indicates 2, all elements of the quantization matrix used for the 64×64 sub-block using zero-out are coded instead of the 8×8 base quantization matrix. This makes it possible to decode a bitstream that selectively realizes quantization matrix code amount reduction and unique quantization control for sub-blocks using zero-out.
[0120] In this embodiment, the subblocks processed using zero-out are only 64×64, but the subblocks processed using zero-out are not limited to this. For example, among the orthogonal transform coefficients corresponding to the 32×64 or 64×32 subblocks shown in FIG. 7(c) or FIG. 7(b), the lower half or right half of the 32×32 orthogonal transform coefficients may not be decoded, and only the upper half or left half of the quantization coefficients may be decoded. In this case, only the upper half or left half of the 32×32 orthogonal transform coefficients are subject to decoding and inverse quantization, and the upper half or left half of the 32×32 orthogonal transform coefficients are quantized using a quantization matrix different from that in FIG. 8(b).
[0121] Furthermore, the value of the quantization matrix corresponding to the DC coefficient located at the upper left corner among the generated orthogonal transform coefficients, which is considered to have the greatest effect on image quality, may be decoded and set separately from the values of each element of the 8×8 base matrix. Figures 12(b) and 12(c) show an example in which the value of the element located at the upper left corner, which corresponds to the DC component, is changed compared to Figures 8(b) and 8(c). In this case, the quantization matrix shown in Figures 12(b) and 12(c) can be set by separately decoding information indicating "2" located in the DC part in addition to the information of the base quantization matrix in Figure 8(a). This makes it possible to decode a bitstream in which more detailed quantization control has been performed on the DC component of the orthogonal transform coefficient that has the greatest effect on image quality.
[0122] <Embodiment 3> In the above embodiment, each processing unit shown in Fig. 1 and Fig. 2 is described as being configured by hardware. However, the processing performed by each processing unit shown in these figures may be configured by a computer program.
[0123] FIG. 5 is a block diagram showing an example of the hardware configuration of a computer applicable to the image display device according to each of the above embodiments.
[0124] The CPU 501 controls the entire computer using computer programs and data stored in the RAM 502 and ROM 503, and executes the processes described above as being performed by the image processing device according to each of the above embodiments. That is, the CPU 501 functions as each of the processing units shown in FIGS.
[0125] The RAM 502 has an area for temporarily storing computer programs and data loaded from an external storage device 506, data acquired from the outside via an I / F (interface) 507, etc. Furthermore, the RAM 502 has a work area used when the CPU 501 executes various processes. That is, the RAM 502 can be allocated as a frame memory, for example, or can provide various other areas as appropriate.
[0126] The ROM 503 stores setting data for the computer, a boot program, etc. The operation unit 504 is composed of a keyboard, a mouse, etc., and can be operated by a user of the computer to input various instructions to the CPU 501. The display unit 505 displays the results of processing by the CPU 501. The display unit 505 is composed of, for example, a liquid crystal display.
[0127] The external storage device 506 is a large-capacity information storage device, such as a hard disk drive. The external storage device 506 stores an operating system (OS) and computer programs for causing the CPU 501 to realize the functions of the various units shown in Figures 1 and 2. Furthermore, the external storage device 506 may store various image data to be processed.
[0128] Computer programs and data stored in the external storage device 506 are loaded into the RAM 502 as appropriate under the control of the CPU 501, and become the subject of processing by the CPU 501. A network such as a LAN or the Internet, and other devices such as a projector or display device can be connected to the I / F 507, and the computer can obtain and transmit various information via this I / F 507. Reference numeral 508 denotes a bus that connects the above-mentioned components.
[0129] The operation of the above-mentioned configuration is controlled mainly by CPU 501 as explained in the above-mentioned flow chart.
[0130] (Other Examples) Each embodiment can also be achieved by supplying a storage medium on which a computer program code that realizes the above-mentioned functions is recorded to a system, and the system reads and executes the computer program code. In this case, the computer program code read from the storage medium itself realizes the functions of the above-mentioned embodiments, and the storage medium on which the computer program code is stored constitutes the present invention. Also included is a case in which an operating system (OS) running on a computer performs part or all of the actual processing based on the instructions of the program code, and the above-mentioned functions are realized by that processing.
[0131] Furthermore, the present invention may be realized in the following form: That is, the computer program code read from the storage medium is written into a memory provided in a function expansion card inserted into a computer or a function expansion unit connected to the computer. Then, based on the instructions of the computer program code, a CPU provided in the function expansion card or function expansion unit performs part or all of the actual processing to realize the above-mentioned functions.
[0132] When the present invention is applied to the above-mentioned storage medium, the storage medium stores computer program code corresponding to the flowcharts described above. [Explanation of symbols]
[0133] 101, 112, 201, 208 terminals 102 Block division section 103 Quantization matrix storage unit 104 Prediction Department 105 Transformation and Quantization Section 106, 204 Inverse quantization and inverse transformation section 107, 205 Image playback unit 108, 206 Frame Memory 109, 207 In-loop filter section 110 Encoding section 111 Integrated coding unit 113 Quantization matrix coding unit 202 Separate Decoding Unit 203 Decoding section 209 Quantization matrix decoding unit
Claims
1. 1. An image decoding device capable of decoding an image from a bitstream using a plurality of blocks including a first block of P×Q pixels (P and Q are integers) and a second block of N×M pixels (N is an integer satisfying N<P and M is an integer satisfying M<Q), a decoding means for decoding data corresponding to a first set of quantized transform coefficients corresponding to the first block and data corresponding to a second set of quantized transform coefficients corresponding to the second block from the bit stream; an inverse quantization means for deriving a first set of transform coefficients representing frequency components from the first set of quantized transform coefficients using a first quantization matrix having N×M elements, and for deriving a second set of transform coefficients representing frequency components from the second set of quantized transform coefficients using a second quantization matrix having N×M elements; an inverse transform means for deriving a first set of prediction errors corresponding to the first block by performing an inverse transform process on the first set of transform coefficients, and for deriving a second set of prediction errors corresponding to the second block by performing an inverse transform process on the second set of transform coefficients; a reproduction unit that reproduces at least image data based on the first group of prediction errors and predicted image data; having The predicted image data can be derived using a prediction method that combines intra prediction and inter prediction, When the block to be decoded is the first block, the inverse transform means multiplies the first transform coefficient group, which is N×M transform coefficients, by an M×Q matrix to derive N×Q intermediate values, and further multiplies the N×Q intermediate values by a P×N matrix to derive the first prediction error group, which is P×Q prediction errors, from the first transform coefficient group; the first quantization matrix having the N×M elements is a quantization matrix that includes some elements in a third quantization matrix having R×S elements (R is an integer satisfying R≦N and S is an integer satisfying S≦M) but does not include other elements in the third quantization matrix, the second quantization matrix having the N×M elements is a quantization matrix that includes all elements in a fourth quantization matrix having R×S elements; the third quantization matrix is different from the fourth quantization matrix; the first quantization matrix is a quantization matrix that is configured with the part of elements in the third quantization matrix except for an element corresponding to a DC component, the second quantization matrix is a quantization matrix that is composed of all of the elements in the fourth quantization matrix except for an element corresponding to a DC component; The inverse quantization means further uses a quantization parameter to derive the set of transform coefficients.
2. An image decoding device comprising:
2. 1. An image decoding method capable of decoding an image from a bitstream using a plurality of blocks including a first block of P×Q pixels (P and Q are integers) and a second block of N×M pixels (N is an integer satisfying N<P and M is an integer satisfying M<Q), a decoding step of decoding data corresponding to a first set of quantized transform coefficients corresponding to the first block and data corresponding to a second set of quantized transform coefficients corresponding to the second block from the bitstream; a dequantization step of deriving a first set of transform coefficients representing frequency components from the first set of quantized transform coefficients using a first quantization matrix having N×M elements, and deriving a second set of transform coefficients representing frequency components from the second set of quantized transform coefficients using a second quantization matrix having N×M elements; an inverse transform step of deriving a first set of prediction errors corresponding to the first block by performing an inverse transform process on the first set of transform coefficients, and deriving a second set of prediction errors corresponding to the second block by performing an inverse transform process on the second set of transform coefficients; a reproduction step of at least reproducing image data based on the first set of prediction errors and predicted image data; having The predicted image data can be derived using a prediction method that combines intra prediction and inter prediction, When the block to be decoded is the first block, in the inverse transform step, the first transform coefficient group, which is N×M transform coefficients, is multiplied by an M×Q matrix to derive N×Q intermediate values, and further, the first prediction error group, which is P×Q prediction errors, is derive from the first transform coefficient group by multiplying the N×Q intermediate values by a P×N matrix; the first quantization matrix having the N×M elements is a quantization matrix that includes some elements in a third quantization matrix having R×S elements (R is an integer satisfying R≦N and S is an integer satisfying S≦M) but does not include other elements in the third quantization matrix, the second quantization matrix having the N×M elements is a quantization matrix that includes all elements in a fourth quantization matrix having R×S elements; the third quantization matrix is different from the fourth quantization matrix; the first quantization matrix is a quantization matrix that is configured with the part of elements in the third quantization matrix except for an element corresponding to a DC component, the second quantization matrix is a quantization matrix that is composed of all of the elements in the fourth quantization matrix except for an element corresponding to a DC component; The quantization parameter is further used in the inverse quantization step to derive the set of transform coefficients.
2. An image decoding method comprising:
3. The first and second blocks are square blocks.
3. The image decoding method according to claim 2.
4. The P and the Q are 64, and the N and the M are 32.
3. The image decoding method according to claim 2.
5. The P and the Q are 128, and the N and the M are 32.
3. The image decoding method according to claim 2.
6. the first set of transform coefficients being N×M transform coefficients; The second set of transform coefficients is N×M transform coefficients.
3. The image decoding method according to claim 2.
7. The first and second blocks are non-square blocks.
3. The image decoding method according to claim 2.
8. the first set of prediction errors is P×Q prediction errors; The second set of prediction errors is N×M prediction errors.
3. The image decoding method according to claim 2.
9. 1. An image encoding device capable of encoding an image using a plurality of blocks including a first block of P×Q pixels (P and Q are integers) and a second block of N×M pixels (N is an integer satisfying N<P and M is an integer satisfying M<Q), a transform means for deriving a first set of transform coefficients by performing a transform process on a first set of prediction errors corresponding to the first block, and for deriving a second set of transform coefficients by performing a transform process on a second set of prediction errors corresponding to the second block; a quantization means for quantizing the first set of transform coefficients using a first quantization matrix having N×M elements to derive a first set of quantized transform coefficients, and for quantizing the second set of transform coefficients using a second quantization matrix having N×M elements to derive a second set of quantized transform coefficients; encoding means for encoding data corresponding to the first set of quantized transform coefficients corresponding to the first block and data corresponding to the second set of quantized transform coefficients corresponding to the second block; having At least the first set of prediction errors can be derived using a combined intra- and inter-prediction prediction method, When the block to be coded is the first block, the transform means multiplies the first prediction error group, which is P×Q prediction errors, by a Q×M matrix to derive P×M intermediate values, and further multiplies the P×M intermediate values by an N×P matrix to derive the first transform coefficient group, which is N×M transform coefficients, from the first prediction error group; the first quantization matrix having the N×M elements is a quantization matrix that includes some elements in a third quantization matrix having R×S elements (R is an integer satisfying R≦N and S is an integer satisfying S≦M) but does not include other elements in the third quantization matrix, the second quantization matrix having the N×M elements is a quantization matrix that includes all elements in a fourth quantization matrix having R×S elements; the third quantization matrix is different from the fourth quantization matrix; the first quantization matrix is a quantization matrix that is configured with the part of elements in the third quantization matrix except for an element corresponding to a DC component, the second quantization matrix is a quantization matrix that is composed of all of the elements in the fourth quantization matrix except for an element corresponding to a DC component; The quantization means further uses a quantization parameter to derive the set of quantized transform coefficients.
1. An image encoding device comprising:
10. 1. An image coding method capable of coding an image using a plurality of blocks including a first block of P×Q pixels (P and Q are integers) and a second block of N×M pixels (N is an integer satisfying N<P and M is an integer satisfying M<Q), a transform step of deriving a first set of transform coefficients by performing a transform process on a first set of prediction errors corresponding to the first block, and deriving a second set of transform coefficients by performing a transform process on a second set of prediction errors corresponding to the second block; a quantization step of quantizing the first set of transform coefficients using a first quantization matrix having N×M elements to derive a first set of quantized transform coefficients, and quantizing the second set of transform coefficients using a second quantization matrix having N×M elements to derive a second set of quantized transform coefficients; an encoding step of encoding data corresponding to the first set of quantized transform coefficients corresponding to the first block and data corresponding to the second set of quantized transform coefficients corresponding to the second block; having At least the first set of prediction errors can be derived using a combined intra- and inter-prediction prediction method, When the block to be coded is the first block, in the transform step, the first prediction error group, which is P×Q prediction errors, is multiplied by a Q×M matrix to derive P×M intermediate values, and further, the first transform coefficient group, which is N×M transform coefficients, is derive from the first prediction error group by multiplying the P×M intermediate values by an N×P matrix; the first quantization matrix having the N×M elements is a quantization matrix that includes some elements in a third quantization matrix having R×S elements (R is an integer satisfying R≦N and S is an integer satisfying S≦M) but does not include other elements in the third quantization matrix, the second quantization matrix having the N×M elements is a quantization matrix that includes all elements in a fourth quantization matrix having R×S elements; the third quantization matrix is different from the fourth quantization matrix; the first quantization matrix is a quantization matrix that is configured with the part of elements in the third quantization matrix except for an element corresponding to a DC component, the second quantization matrix is a quantization matrix that is composed of all of the elements in the fourth quantization matrix except for an element corresponding to a DC component; A quantization parameter is further used in the quantization step to derive a set of quantized transform coefficients.
13. An image coding method comprising:
11. A program for causing a computer to execute the image decoding method according to any one of claims 2 to 8.
12. A program for causing a computer to execute the image coding method according to claim 10.
Citation Information
Patent Citations
Image encoding device, image encoding method, and program
JP7651615B2
Image processing device and method
WO2018008387A1
Image processing device and image processing method
WO2019188097A1
Image encoder, image encoding method, program, image decoder, image decoding method and program
JP2013038758A