Image decoding device, image decoding method, image encoding device, image encoding method, and computer program
The image decoding apparatus optimizes VVC bitstream efficiency by utilizing flags and positional information to reduce redundant syntax, addressing the issue of unnecessary code expansion in VVC encoding.
Patent Information
- Application Number
- JP2025086645
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2039-06-21
AI Technical Summary
In the VVC coding method, the number of entry_point_offset_minus1 indicating the start position of basic block rows within a slice is redundantly encoded, leading to unnecessary code expansion in the bitstream.
An image decoding apparatus that decodes an image from a bitstream by utilizing a first flag to identify rectangular regions within a slice, along with second and third information to specify the start position of block rows, thereby reducing redundant syntax and optimizing the bitstream's code efficiency.
The proposed solution effectively reduces the amount of code in the bitstream by eliminating redundant syntax elements, enhancing encoding efficiency.
Smart Images

Figure 2025109990000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to image encoding / decoding technology.
Background Art
[0002] As an encoding method for compressed recording of moving images, the HEVC (High Efficiency Video Coding) encoding method (hereinafter referred to as HEVC) is known. In HEVC, in order to improve the encoding efficiency, a basic block having a size larger than that of a conventional macroblock (16×16 pixels) is adopted. This basic block having a large size is called a CTU (Coding Tree Unit), and its size is up to 64×64 pixels. The CTU is further divided into sub-blocks that are units for performing prediction and conversion.
[0003] Also in HEVC, it is possible to divide a picture into a plurality of tiles or slices and perform encoding. There is little data dependency between each tile or slice, and parallel encoding / decoding processing can be performed. One of the major advantages of tile and slice division is that parallel processing can be executed using a multi-core CPU or the like to shorten the processing time.
[0004] Also, each slice is encoded by the conventional binary arithmetic coding method adopted in HEVC. That is, each syntax element is binarized to generate a binary signal. An occurrence probability is given in advance to each syntax element as a table (hereinafter referred to as an occurrence probability table), and the binary signal is arithmetic-coded based on the occurrence probability table. This occurrence probability table is used as decoding information for decoding subsequent codes during decoding. During encoding, it is used as encoding information for subsequent encoding. Then, every time encoding is performed, the occurrence probability table is updated based on statistical information as to whether the encoded binary signal is a symbol with a higher occurrence probability.
[0005] In addition, HEVC has a technique called Wavefront Parallel Processing (hereinafter referred to as WPP) for processing entropy encoding and decoding in parallel. In WPP, by applying the probability table at the time of encoding the block at a specified position in advance to the block at the left end of the next row, it is possible to suppress the decrease in encoding efficiency and perform parallel encoding processing of blocks in row units. In order to enable parallel processing in block row units, the entry_point_offset_minus1 indicating the start position of each block row in the bitstream and the num_entry_point_offsets indicating the number thereof are encoded in the slice header. Patent Document 1 discloses a technique related to WPP.
[0006] In recent years, activities have been started to internationalize an even more efficient coding method as a successor to HEVC. The JVET (Joint Video Experts Team) was established between ISO / IEC and ITU-T, and standardization is underway for the VVC (Versatile Video Coding) coding method (hereinafter referred to as VVC). In VVC, it is being considered to divide a tile into rectangles (bricks) composed of a plurality of block rows. And a slice is configured to include one or more bricks.
Prior Art Documents
Patent Documents
[0007]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0008] In VVC, the blocks that make up a slice can be derived in advance, and furthermore, the number of basic block rows included in the blocks can also be derived from other syntax. Therefore, it is possible to derive the number of entry_point_offset_minus1 indicating the start position of the basic block rows belonging to the slice without using num_entry_point_offset. Therefore, num_entry_point_offset is redundant syntax. The present invention provides a technique for reducing the amount of code in a bitstream by reducing redundant syntax.
Means for Solving the Problem
[0009] One aspect of the present invention is an image decoding apparatus that decodes an image from a bitstream obtained by encoding an image including a rectangular region including one or more block rows composed of a plurality of blocks, a first flag related to enabling parallel processing, and a first flag used to identify the first rectangular region to be processed among the plurality of rectangular regions included in a slice in the image. Information, second information used to identify the last rectangular region to be processed among the plurality of rectangular regions, and third information corresponding to the number of blocks in the vertical direction of the rectangular region in the image, and decoding means for decoding from the bitstream, and the value of the first flag is 1, and a second flag that is a flag decoded from a picture parameter set of the bitstream and is related to the mode of the slice. When the mode in which the slice is rectangular is used, and in a state where a plurality of rectangular regions are included in the horizontal or vertical direction in the target slice included in the image, based on the first information, the second information, and the third information, for the target slice, specifying means for specifying the number of information for specifying the start position of the coded data of the block row, and the decoding means is based on at least the number of information for specifying the start position specified by the specifying means and the information for specifying the start position. Characterized by decoding the coded data of the block row.
Effect of the Invention
[0010] According to the configuration of the present invention, the amount of code of the bit stream can be reduced by reducing redundant syntax.
Brief Description of Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Embodiments for Carrying Out the Invention
[0012] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the invention according to the claims. Although a plurality of features are described in the embodiments, not all of these plurality of features are essential to the invention, and the plurality of features may be arbitrarily combined. Further, in the accompanying drawings, the same or similar configurations are denoted by the same reference numerals, and redundant descriptions are omitted.
[0013] [First Embodiment] First, a functional configuration example of the image encoding apparatus according to the present embodiment will be described with reference to the block diagram of FIG. 1. An input image to be encoded is input to the image division unit 102. The input image may be an image of each frame constituting a moving image or a still image. The image division unit 102 divides the input image into "one or a plurality of tiles". A tile is a set of consecutive basic blocks that cover a rectangular area within the input image. The image division unit 102 further divides each tile into one or a plurality of blocks. A block is a rectangular area (a rectangular area including one or more block rows of a plurality of blocks having a size equal to or smaller than that of the tile) composed of one or a plurality of rows of basic blocks (basic block rows) within the tile. The image division unit 102 further divides the input image into slices composed of "one or a plurality of tiles" or "one or more blocks within one tile". A slice is a basic unit of encoding, and header information such as information indicating the type of slice is added for each slice. An example of dividing the input image into 4 tiles, 4 slices, and 11 blocks is shown in FIG. 7. The upper left tile is divided into 1 block, the lower left tile is divided into 2 blocks, the upper right tile is divided into 5 blocks, and the lower right tile is divided into 3 blocks, respectively. And the left slice is configured to include 3 blocks, the upper right slice is configured to include 2 blocks, the middle right slice is configured to include 3 blocks, and the lower right slice is configured to include 3 blocks. The image division unit 102 outputs information regarding the size as division information for each of the tiles, blocks, and slices divided in this way.
[0014] The block division unit 103 divides the image of the basic block row (basic block row image) output from the image division unit 102 into a plurality of basic blocks, and outputs the image in units of basic blocks (block image) to the subsequent stage.
[0015] The prediction unit 104 divides the image in basic block units into sub-blocks, performs intra prediction which is in-frame prediction in sub-block units, inter prediction which is inter-frame prediction, etc., and generates a predicted image. Intra prediction across blocks (intra prediction using pixels of blocks in other blocks) and prediction of motion vectors across blocks (prediction of motion vectors using motion vectors of blocks in other blocks) are not performed. Further, the prediction unit 104 calculates and outputs a prediction error from the input image and the predicted image. Also, the prediction unit 104 outputs information necessary for prediction (prediction information), such as sub-block division method, prediction mode, motion vector, etc., together with the prediction error.
[0016] The transform and quantization unit 105 orthogonally transforms the prediction error in sub-block units to obtain transform coefficients, and quantizes the obtained transform coefficients to obtain quantization coefficients. The inverse quantization and inverse transform unit 106 inverse quantizes the quantization coefficients output from the transform and quantization unit 105 to reproduce the transform coefficients, and further inverse orthogonally transforms the reproduced transform coefficients to reproduce the prediction error.
[0017] The frame memory 108 functions as a memory for storing the reproduced image. The image reproduction unit 107 appropriately refers to the frame memory 108 based on the prediction information output from the prediction unit 104 to generate a predicted image, and generates and outputs a reproduced image from the predicted image and the input prediction error.
[0018] The in-loop filter unit 109 performs in-loop filter processing such as deblocking filter and sample adaptive offset on the reproduced image, and outputs the image (filtered image) subjected to the in-loop filter processing.
[0019] The encoding unit 110 generates coded data (encoded data) by encoding the quantization coefficients output from the transform and quantization unit 105 and the prediction information output from the prediction unit 104, and outputs the generated coded data.
[0020] The integrated encoding unit 111 generates header encoded data using the segmentation information output from the image segmentation unit 102, and generates and outputs a bit stream including the generated header encoded data and the encoded data output from the encoding unit 110. The control unit 199 controls the operation of the entire image encoding device, and controls the operation of each functional unit of the above-described image encoding device.
[0021] Next, the encoding process for the input image by the image encoding device having the configuration shown in FIG. 1 will be described. In the present embodiment, for ease of explanation, only the process of intra prediction encoding will be described, but the present invention is not limited thereto and is also applicable to the process of inter prediction encoding. Further, in the present embodiment, for the purpose of specific explanation, the block segmentation unit 103 will be described as dividing the basic block row image output from the image segmentation unit 102 in units of "basic blocks having a size of 64×64 pixels".
[0022] The image segmentation unit 102 divides the input image into tiles and blocks. An example of dividing the input image by the image segmentation unit 102 is shown in FIG. 8. In the present embodiment, as shown in FIG. 8(a), an input image having a size of 1152×1152 pixels is divided into nine tiles (the size of one tile is 384×384 pixels). Each tile is assigned an ID (tile ID) in raster order from the upper left, and the tile ID of the upper left tile is 0, and the tile ID of the lower right tile is 8.
[0023] In addition, an example of dividing a tile, a brick, and a slice in the input image is shown in FIG. 8(b). As shown in FIG. 8(b), the tile with tile ID = 0 and the tile with tile ID = 7 are each divided into two bricks (the size of each brick is 384×192 pixels). The tile with tile ID = 2 is divided into two bricks (the size of the upper brick is 384×128 pixels and the size of the lower brick is 384×256 pixels). The tile with tile ID = 3 is divided into three bricks (the size of each brick is 384×128 pixels). The tiles with tile ID = 1, 4, 5, 6, 8 are not divided into bricks (equivalent to dividing one tile into one brick), and as a result, tile = brick. Each brick is assigned an ID in order from the top in the raster-ordered tiles. The BID shown in FIG. 8(b) is the ID of the brick. The input image is also divided into slices including bricks corresponding to BID = 0 to 2, a slice including the brick corresponding to BID = 3, a slice including the brick corresponding to BID = 4, slices including bricks corresponding to BID = 5 to 8 and 10 to 12, and slices including bricks corresponding to BID = 9 and 13. Note that each slice is also assigned an ID in order from the top in the raster-ordered slices. For example, slice 0 refers to the slice with ID = 0, and slice 4 refers to the slice with ID = 4.
[0024] Then, for each of the divided tiles, bricks, and slices, the image dividing unit 102 outputs information regarding the size as division information to the integrated encoding unit 111. The image dividing unit 102 also divides each brick into basic block row images and outputs the divided basic block row images to the block dividing unit 103.
[0025] The block dividing unit 103 divides the basic block row image output from the image dividing unit 102 into a plurality of basic blocks and outputs block images (64×64 pixels), which are images in basic block units, to the subsequent prediction unit 104.
[0026] The prediction unit 104 divides the image in basic block units into sub-blocks, determines an intra prediction mode such as horizontal prediction or vertical prediction for each sub-block unit, and generates a prediction image from the determined intra prediction mode and the encoded pixels. Further, the prediction unit 104 calculates a prediction error from the input image and the prediction image, and outputs the calculated prediction error to the transform and quantization unit 105. Also, the prediction unit 104 outputs information such as the sub-block division method and the intra prediction mode as prediction information to the encoding unit 110 and the image reproduction unit 107.
[0027] The transform and quantization unit 105 performs an orthogonal transform (orthogonal transform process corresponding to the size of the sub-block) on the prediction error output from the prediction unit 104 in sub-block units to obtain transform coefficients (orthogonal transform coefficients). Then, the transform and quantization unit 105 quantizes the obtained transform coefficients to obtain quantization coefficients. Then, the transform and quantization unit 105 outputs the obtained quantization coefficients to the encoding unit 110 and the inverse quantization and inverse transform unit 106.
[0028] The inverse quantization and inverse transform unit 106 inverse quantizes the quantization coefficients output from the transform and quantization unit 105 to reproduce the transform coefficients, and further inverse orthogonally transforms the reproduced transform coefficients to reproduce the prediction error. Then, the inverse quantization and inverse transform unit 106 outputs the reproduced prediction error to the image reproduction unit 107.
[0029] The image reproduction unit 107 appropriately refers to the frame memory 108 based on the prediction information output from the prediction unit 104 to generate a prediction image, and generates a reproduced image from the prediction image and the prediction error input from the inverse quantization and inverse transform unit 106. Then, the image reproduction unit 107 stores the generated reproduced image in the frame memory 108.
[0030] The in-loop filter unit 109 reads out the reproduced image from the frame memory 108, and performs in-loop filter processing such as a deblocking filter and a sample adaptive offset on the read reproduced image. Then, the in-loop filter unit 109 stores (re-stores) the image subjected to the in-loop filter processing in the frame memory 108 again.
[0031] The symbolization unit 110 generates coded data by entropy-coding the quantized coefficients output from the conversion / quantization unit 105 and the prediction information output from the prediction unit 104. Although the method of entropy-coding is not particularly specified, Golomb coding, arithmetic coding, Huffman coding, etc. can be used. Then, the coding unit 110 outputs the generated coded data to the integrated coding unit 111.
[0032] The integrated coding unit 111 generates header coded data using the segmentation information output from the image segmentation unit 102, and generates and outputs a bitstream by multiplexing the generated header coded data and the coded data output from the coding unit 110. The output destination of the bitstream is not limited to a specific output destination, and it may be output (stored) to an internal or external memory of the image coding device, or transmitted to an external device capable of communicating with the image coding device via a network such as a LAN or the Internet.
[0033] Next, an example of the format of the bitstream (coded data by VVC coded by the image coding device) output by the integrated coding unit 111 is shown in FIG. 6. The bitstream in FIG. 6 includes a sequence parameter set (SPS), which is header information including information related to the coding of the sequence. The bitstream in FIG. 6 also includes a picture parameter set (PPS), which is header information including information related to the coding of the picture. The bitstream in FIG. 6 further includes a slice header (SLH), which is header information including information related to the coding of the slice. The bitstream in FIG. 6 also includes the coded data of each block (blocks 0 to (N-1) in FIG. 6).
[0034] The SPS contains image size information and basic block data partitioning information. The PPS contains tile data partitioning information, which is tile partitioning information, block data partitioning information, which is block partitioning information, slice data partitioning information 0, which is slice partitioning information, and basic block row data synchronization information. The SLH contains slice data partitioning information 1 and basic block row data position information.
[0035] First, the SPS will be described. The SPS contains pic_width_in_luma_samples, which is information 601, as image size information, and pic_height_in_luma_samples, which is information 602. pic_width_in_luma_samples represents the horizontal size (number of pixels) of the input image, and pic_height_in_luma_samples represents the vertical size (number of pixels) of the input image. In this embodiment, since the input image in FIG. 8 is used as the input image, pic_width_in_luma_samples = 1152 and pic_height_in_luma_samples = 1152. The SPS also contains log2_ctu_size_minus2, which is information 603, as basic block data partitioning information. log2_ctu_size_minus2 represents the size of the basic block. The number of pixels in the vertical and horizontal directions of the basic block is indicated by 1<<(log2_ctu_size_minus2 + 2). In this embodiment, since the size of the basic block is 64×64 pixels, the value of log2_ctu_size_minus2 is 4.
[0036] Next, PPS will be described. PPS includes information 604 to 607 as tile data division information. Information 604 is a single_tile_in_pic_flag indicating whether the input image is divided into a plurality of tiles and encoded. When single_tile_in_pic_flag = 1, it indicates that the input image is not divided into a plurality of tiles and encoded. On the other hand, when single_tile_in_pic_flag = 0, it indicates that the input image is divided into a plurality of tiles and encoded.
[0037] Information 605 is information included in the tile data division information when single_tile_in_pic_flag = 0. Information 605 is a uniform_tile_spacing_flag indicating whether each tile has the same size. When uniform_tile_spacing_flag = 1, it indicates that each tile has the same size, and when uniform_tile_spacing_flag = 0, it indicates that there are tiles with different sizes.
[0038] Information 606 and Information 607 are information included in tile data division information when uniform_tile_spacing_flag = 1. Information 606 is tile_cols_width_minus1 indicating (the number of basic blocks in the horizontal direction of the tile - 1). Information 607 is tile_rows_height_minus1 indicating (the number of basic blocks in the vertical direction of the tile - 1). The number of tiles in the horizontal direction of the input image is obtained as the quotient when the number of basic blocks in the horizontal direction of the input image is divided by the number of basic blocks in the horizontal direction of the tile. If a remainder occurs in this division, the number obtained by adding 1 to the quotient is taken as "the number of tiles in the horizontal direction of the input image". Also, the number of tiles in the vertical direction of the input image is obtained as the quotient when the number of basic blocks in the vertical direction of the input image is divided by the number of basic blocks in the vertical direction of the tile. If a remainder occurs in this division, the number obtained by adding 1 to the quotient is taken as "the number of tiles in the vertical direction of the input image". Also, the total number of tiles in the input image can be obtained by calculating the number of tiles in the horizontal direction of the input image × the number of tiles in the vertical direction of the input image.
[0039] Note that when uniform_tile_spacing_flag = 0, since tiles of different sizes from others are included, the number of tiles in the horizontal direction of the input image, the number of tiles in the vertical direction of the input image, and the vertical and horizontal sizes of each tile are coded.
[0040] Also, the PPS includes Information 608 to 613 as brick data division information. Information 608 is brick_splitting_present_flag. When brick_splitting_present_flag = 1, it indicates that one or more tiles in the input image are divided into multiple bricks. On the other hand, when brick_splitting_present_flag = 0, it indicates that each tile in the input image is composed of a single brick.
[0041] Information 609 is information included in brick data division information when brick_splitting_present_flag = 1. Information 609 is brick_split_flag[] indicating whether each tile is divided into multiple bricks. brick_split_flag[] indicating whether the i-th tile is divided into multiple bricks is denoted as brick_split_flag[i]. When brick_split_flag[i] = 1, it indicates that the i-th tile is divided into multiple bricks, and when brick_split_flag[i] = 0, it indicates that the i-th tile is composed of a single brick.
[0042] Information 610 is uniform_brick_spacing_flag[i] indicating whether the sizes of all bricks constituting the i-th tile are the same when brick_split_flag[i] = 1. If brick_split_flag[i] = 0 for all i, information 610 is not included in the brick data division information. Information 610 includes uniform_brick_spacing_flag[i] for i satisfying brick_split_flag[i] = 1. When uniform_brick_spacing_flag[i] = 1, it indicates that the sizes of all bricks constituting the i-th tile are the same. On the other hand, when uniform_brick_spacing_flag[i] = 0, it indicates that there is a brick with a size different from others among the bricks constituting the i-th tile.
[0043] Information 611 is information included in the brick data division information when uniform_brick_spacing_flag[i] = 1. Information 611 is brick_height_minus1[i] indicating (the number of basic blocks in the vertical direction of the bricks in the i-th tile - 1).
[0044] Note that the number of basic blocks in the vertical direction of a brick can be obtained by dividing the number of pixels in the vertical direction of the brick by the number of pixels in the vertical direction of a basic block (64 pixels in this embodiment). Also, the number of bricks constituting a tile is obtained as the quotient when the number of basic blocks in the vertical direction of the tile is divided by the number of basic blocks in the vertical direction of a brick. If a remainder occurs in this division, the number obtained by adding 1 to the quotient is defined as the "number of bricks constituting the tile". For example, assume that the number of basic blocks in the vertical direction of a tile is 10 and the value of brick_height_minus1 is 2. At this time, this tile is divided into four bricks in order from the top: a brick with 3 rows of basic blocks, a brick with 3 rows of basic blocks, a brick with 3 rows of basic blocks, and a brick with 1 row of basic blocks.
[0045] The information 612 is num_brick_rows_minus1[i] which indicates (the number of bricks constituting the i-th tile - 1) for i that satisfies uniform_brick_spacing_flag[i]=0.
[0046] Note that in this embodiment, when uniform_brick_spacing_flag[i]=0, num_brick_rows_minus1[i] which indicates (the number of bricks constituting the i-th tile - 1) is included in the brick data division information. However, the present invention is not limited to this.
[0047] For example, assume that when brick_split_flag[i] = 1, the number of bricks constituting the i-th tile is 2 or more. Then, num_brick_rows_minus2[i] indicating (the number of bricks constituting the tile - 2) may be encoded instead of num_brick_rows_minus1[i]. By doing so, the number of bits of the syntax indicating the number of bricks constituting the tile can be reduced. For example, when the tile is composed of 2 bricks and num_brick_rows_minus1[i] is Golomb-encoded, 3-bit data of "010" indicating "1" is encoded. On the other hand, when num_brick_rows_minus2[i] indicating (the number of bricks constituting the tile - 2) is Golomb-encoded, 1-bit data of "0" indicating 0 is encoded.
[0048] The information 613 is brick_row_height_minus1[i][j] indicating (the number of basic blocks in the vertical direction of the j-th brick in the i-th tile - 1) for i satisfying uniform_brick_spacing_flag[i] = 0. brick_row_height_minus1[i][j] is encoded by the number of num_brick_rows_minus1[i]. Note that when the above-mentioned num_brick_rows_minus2[i] is used, brick_row_height_minus1[i][j] is encoded by the number of num_brick_rows_minus2[i] + 1. The number of basic blocks in the vertical direction of the bottom brick in the tile can be obtained by subtracting the sum of "brick_row_height_minus1 + 1" from the number of basic blocks in the vertical direction of the tile. For example, assume that the number of basic blocks in the vertical direction of the tile = 10, num_brick_rows_minus1 = 3, brick_row_height_minus1 = 2, 1, 2. At this time, the number of basic blocks in the vertical direction of the bottom brick in the tile is 10 - (3 + 2 + 3) = 2.
[0049] PPS also contains information 614 to 618 as slice data division information 0. Information 614 is single_brick_per_slice_flag. When single_brick_per_slice_flag = 1, it indicates that all slices in the input image are composed of a single brick. On the other hand, when single_brick_per_slice_flag = 0, it indicates that one or more slices in the input image are composed of multiple bricks. That is, it indicates that each slice is composed of only one brick.
[0050] Information 615 is rect_slice_flag, which is the information included in slice data division information 0 when single_brick_per_slice_flag = 0. rect_slice_flag indicates whether the tiles included in the slice are in raster order or rectangular. Figure 9(a) shows the relationship between tiles and slices when rect_slice_flag = 0, indicating that the tiles in the slice are encoded in raster order. On the other hand, Figure 9(b) shows the relationship between tiles and slices when rect_slice_flag = 1, indicating that multiple tiles in the slice are rectangular.
[0051] Information 616 is num_slices_in_pic_minus1, which is the information included in slice data division information 0 when rect_slice_flag = 1 and single_brick_per_slice_flag = 0. num_slices_in_pic_minus1 indicates (the number of slices in the input image - 1).
[0052] Information 617 is top_left_brick_idx[i] which indicates the index of the top - left brick of each slice (the i - th slice) in the input image.
[0053] The information 618 is bottom_right_brick_idx_delta[i] which indicates the difference between the index of the top-left brick and the index of the bottom-right brick in the i-th slice of the input image. Here, the "top-left brick in the i-th slice of the input image" is the brick that is processed first in that slice. Also, the "bottom-right brick in the i-th slice of the input image" is the brick that is processed last in that slice.
[0054] Here, the range of i is from 0 to num_slices_in_pic_minus1. However, since the index of the top-left brick of the first slice in the frame is determined to be 0, top_left_brick_idx[0] of the first slice is not encoded. In this embodiment, the range of i for top_left_brick_idx[i] and bottom_right_brick_idx_delta[i] is set to 0 to num_slices_in_pic_minus1, but it is not limited to this. For example, the bottom-right brick of the last slice (the (num_slices_in_pic_minus1)-th slice) is determined to be the brick with the largest BID. Therefore, bottom_right_brick_idx_delta[num_slices_in_pic_minus1] may not need to be encoded. Furthermore, since the bricks included in the slices other than the last slice are already specified from top_left_brick_idx[i] and bottom_right_brick_idx_delta[i] of the slices other than the last slice. Therefore, the bricks included in the last slice can be specified as all the bricks not included in the previous slices. In that case, the top-left brick of the last slice can be specified as the brick with the smallest BID among the remaining bricks. Therefore, top_left_brick_idx[num_slices_in_pic_minus1] may not need to be encoded either. By doing so, the amount of bits in the header can be further reduced.
[0055] In addition, information 619 is encoded and included in PPS as basic block row data synchronization information. Information 619 is the entropy_coding_sync_enabled_flag. When entropy_coding_sync_enabled_flag = 1, the probability table at the time of processing the basic block at a predetermined position in the basic block row adjacent above is applied to the leftmost block. As a result, parallel processing of entropy encoding / decoding can be performed in units of basic block rows.
[0056] Next, SLH will be described. Information 620 to 621 is encoded and included in SLH as slice data division information 1. Information 620 is the slice_address that is included in the slice data division information 1 when rect_slice_flag = 1 or the number of bricks in the input image is 2 or more. slice_address indicates the BID at the start of the slice when rect_slice_flag = 0, and indicates the number of the current slice when rect_slice_flag = 1.
[0057] Information 621 is the num_bricks_in_slice_minus1 that is included in the slice data division information 1 when rect_slice_flag = 0 and single_brick_per_slice_flag = 0. num_bricks_in_slice_minus1 indicates (the number of bricks in the slice - 1).
[0058] Information 622 is included in SLH as basic block row data position information. Information 622 is entry_point_offset_minus1[]. entry_point_offset_minus1[] is encoded and included in the basic block row data position information by the number of (the number of basic block rows in the slice - 1) when entropy_coding_sync_enabled_flag = 1.
[0059] entry_point_offset_minus1[] represents the entry point of the code data of the basic block line, that is, the starting position of the code data of the basic block line. entry_point_offset_minus1[j - 1] indicates the entry point of the code data of the j-th basic block line. Since the starting position of the code data of the 0-th basic block line is the same as the starting position of the code data of the slice to which the basic block line belongs, it is omitted. And, {the size of the code data of the (j - 1)-th basic block line - 1} is encoded as entry_point_offset_minus1[j - 1].
[0060] Next, the encoding process of the input image by the image encoding device in this embodiment (the generation process of the bitstream having the configuration shown in FIG. 6) will be described according to the flowchart of FIG. 3.
[0061] First, in step S301, the image division unit 102 divides the input image into tiles, bricks, and slices. Then, for each of the divided tiles, bricks, and slices, the image division unit 102 outputs information regarding the size as division information to the integrated encoding unit 111. Also, the image division unit 102 divides each brick into basic block line images and outputs the divided basic block line images to the block division unit 103.
[0062] In step S302, the block division unit 103 divides the basic block line image into a plurality of basic blocks and outputs block images, which are images in basic block units, to the subsequent prediction unit 104.
[0063] In step S303, the prediction unit 104 divides the image in units of basic blocks output from the block division unit 103 into sub-blocks, determines the intra prediction mode in units of sub-blocks, and generates a predicted image from the determined intra prediction mode and the encoded pixels. Further, the prediction unit 104 calculates a prediction error from the input image and the predicted image, and outputs the calculated prediction error to the conversion / quantization unit 105. Also, the prediction unit 104 outputs information such as the sub-block division method and the intra prediction mode as prediction information to the encoding unit 110 and the image reproduction unit 107.
[0064] In step S304, the conversion / quantization unit 105 orthogonally transforms the prediction error output from the prediction unit 104 in units of sub-blocks to obtain conversion coefficients (orthogonal conversion coefficients). Then, the conversion / quantization unit 105 quantizes the obtained conversion coefficients to obtain quantization coefficients. Then, the conversion / quantization unit 105 outputs the obtained quantization coefficients to the encoding unit 110 and the inverse quantization / inverse conversion unit 106.
[0065] In step S305, the inverse quantization / inverse conversion unit 106 inverse quantizes the quantization coefficients output from the conversion / quantization unit 105 to reproduce the conversion coefficients, and further inverse orthogonally transforms the reproduced conversion coefficients to reproduce the prediction error. Then, the inverse quantization / inverse conversion unit 106 outputs the reproduced prediction error to the image reproduction unit 107.
[0066] In step S306, the image reproduction unit 107 appropriately refers to the frame memory 108 based on the prediction information output from the prediction unit 104 to generate a predicted image, and generates a reproduced image from the predicted image and the prediction error input from the inverse quantization / inverse conversion unit 106. Then, the image reproduction unit 107 stores the generated reproduced image in the frame memory 108.
[0067] In step S307, the encoding unit 110 generates coded data by entropy encoding the quantization coefficients output from the conversion / quantization unit 105 and the prediction information output from the prediction unit 104.
[0068] Here, when entropy_coding_sync_enabled_flag = 1, the probability occurrence table at the time of processing the basic block at a predetermined position in the basic block row adjacent above is applied before processing the basic block at the left end of the next basic block row. In the present embodiment, it is described as if entropy_coding_sync_enabled_flag = 1.
[0069] In step S308, the control unit 199 determines whether or not the encoding of all the basic blocks in the slice has been completed. As a result of this determination, if the encoding of all the basic blocks in the slice has been completed, the process proceeds to step S309. On the other hand, if there remains an uncoded basic block (uncoded basic block) among the basic blocks in the slice, the process proceeds to step S303 to encode the uncoded basic block.
[0070] In step S309, the integrated encoding unit 111 generates header encoded data using the division information output from the image division unit 102, and generates and outputs a bit stream including the generated header encoded data and the encoded data output from the encoding unit 110.
[0071] When the input image is divided as shown in FIG. 8, single_tile_in_pic_flag of the tile data division information is 0, and uniform_tile_spacing_flag is 1. Also, tile_cols_width_minus1 is 5, and tile_rows_height_minus1 is 5.
[0072] The brick_splitting_present_flag of the brick data division information is 1. Also, the tiles corresponding to tile IDs = 1, 4, 5, 6, 8 are not divided into bricks. Thus, brick_split_flag[1], brick_split_flag[4], brick_split_flag[5], brick_split_flag[6], brick_split_flag[8] are 0. Also, the tiles corresponding to tile IDs = 0, 2, 3, 7 are divided into bricks. Thus, brick_split_flag[0], brick_split_flag[2], brick_split_flag[3], brick_split_flag[7] are 1.
[0073] Also, the tiles corresponding to tile IDs = 0, 3, 7 are all divided into bricks of the same size. Thus, uniform_brick_spacing_flag[0], uniform_brick_spacing_flag[3], uniform_brick_spacing_flag[7] are all 1. For the tile corresponding to tile ID = 2, the size of the brick with BID = 3 is different from the size of the brick with BID = 4. Thus, uniform_brick_spacing_flag[2] is 0.
[0074] brick_height_minus1[0] is 2, and brick_height_minus1[3] is 1. Also, brick_height_minus1[7] is 2. Note that brick_height_minus1 is encoded when uniform_brick_spacing is 1.
[0075] brick_row_height_minus1[2][0] is 1. Note that when encoding the above syntax of num_brick_rows_minus2[2] instead of num_brick_rows_minus1[2], the value becomes 0.
[0076] Also, in this embodiment, since one slice contains a plurality of bricks, single_brick_per_slice_flag in the slice data division information 0 is 0. Also, in this embodiment, since a slice encloses a plurality of tiles in a rectangle, rect_slice_flag is 1. As shown in FIG. 8(b), since the number of slices in the input image is 5, num_slices_in_pic_minus1 becomes 4.
[0077] Since it is obvious that top_left_brick_idx[0] of slice 0 is 0, it is not encoded, and bottom_right_brick_idx_delta[0] is 2 (= 2 - 0). top_left_brick_idx[1] of slice 1 is 3, and bottom_right_brick_idx_delta[1] is 0 (= 3 - 3). top_left_brick_idx[2] of slice 2 is 4, and bottom_right_brick_idx_delta[2] is 0 (= 4 - 4). top_left_brick_idx[3] of slice 3 is 5, and bottom_right_brick_idx_delta[3] is 7 (= 12 - 5). top_left_brick_idx[4] of slice 4 is 9, and bottom_right_brick_idx_delta[4] is 4 (= 13 - 9). Note that, as described above, top_left_brick_idx[4] and bottom_right_brick_idx_delta[4] do not have to be encoded.
[0078] Also, for entry_point_offset_minus1[], (the size of the coded data of the (j - 1)-th basic block row in the slice - 1) sent from the coding unit 110 is set as entry_point_offset_minus1[j - 1] and coded. The number of entry_point_offset_minus1[] in the slice is equal to (the number of basic block rows in the slice - 1). In this embodiment, since bottom_right_brick_idx_delta[0] is 2, it can be seen that slice 0 consists of bricks with BID = 0 to 2.
[0079] Here, brick_height_minus1[0] is 2 and tile_rows_height_minus1 is 5. Thus, the number of basic block rows of the brick corresponding to BID = 0 and the number of basic block rows of the brick corresponding to BID = 1 are both brick_height_minus1[0]+1 = 3. Also, the number of basic block rows of the brick corresponding to BID = 2 is tile_rows_height_minus1+1 = 6. Therefore, the number of basic block rows in slice 0 is 3 + 3 + 6 = 12. Thus, the range of j is 0 to 10.
[0080] For slice 1, since top_left_brick_idx[1] is 3 and bottom_right_brick_idx_delta[1] is 0, it can be seen that it consists of the brick with BID = 3. Also, since brick_row_height_minus1[2][0] is 1, the number of basic block rows in slice 1 (the brick with BID = 3) is 2. Thus, the range of j is only 0.
[0081] Regarding Slice 2, since top_left_brick_idx[2] is 4 and bottom_right_brick_idx_delta[2] is 0, it can be seen that it consists of the brick with BID = 4. Also, brick_row_height_minus1[2][0] is 1, num_brick_rows_minus1[2] is 1, and tile_rows_height_minus1 is 5. Thus, the number of basic brick rows of Slice 2 (the brick with BID = 4) is {(tile_rows_height_minus1 + 1) - (brick_row_height_minus1[2][0] + 1)} = 4. Therefore, the range of j is 0 to 2.
[0082] Slice 3 is composed of the bricks with BID = 5 to 8, 10 to 12. Since tile_rows_height_minus1 is 5, it can be seen that the number of basic brick rows of the bricks with BID = 8 and 10 is 6 (= tile_rows_height_minus1 + 1). Also, since brick_height_minus1[3] is 1, it can be seen that the number of basic brick rows of the bricks with BID = 5 to 7 is 2 (= brick_height_minus1[3] + 1). Also, since brick_height_minus1[7] is 2, it can be seen that the number of basic brick rows of the bricks with BID = 11 and 12 is 3 (= brick_height_minus1[7] + 1). Thus, it can be seen that the total number of basic brick rows of all the bricks constituting Slice 3 is 2 + 2 + 2 + 6 + 6 + 3 + 3 = 24. Therefore, the range of j is 0 to 22.
[0083] Slice 4 is composed of the brick with BID = 9 and the brick with BID = 13. Since tile_rows_height_minus1 is 5, the number of basic brick rows of each of the brick with BID = 9 and the brick with BID = 13 is 6 (= tile_rows_height_minus1 + 1). Thus, it can be seen that the total number of basic brick rows of all the bricks constituting Slice 4 is 6 + 6 = 12. Therefore, the range of j is 0 to 10.
[0084] Through such processing, the number of basic block lines in each slice is determined. In this embodiment, since the number of entry_point_offset_minus1 can be derived from other syntax, it is not necessary to encode num_entry_point_offset and include it in the header as in the prior art. Therefore, according to this embodiment, the data amount of the bitstream can be reduced.
[0085] In step S310, the control unit 199 determines whether or not the encoding of all the basic blocks in the input image has been completed. As a result of this determination, if the encoding of all the basic blocks in the input image has been completed, the process proceeds to step S311. On the other hand, if there are still basic blocks in the input image that have not been encoded yet, the process proceeds to step S303, and the subsequent processing is performed on the basic blocks that have not been encoded yet.
[0086] In step S311, the in-loop filter unit 109 performs in-loop filter processing on the reconstructed image generated in step S306, and outputs the image subjected to the in-loop filter processing.
[0087] Thus, according to this embodiment, it is not necessary to encode and include in the bitstream information indicating how many pieces of information indicating the start position of the coded data of the basic block lines of the brick are encoded, and a bitstream from which the information can be derived can be generated.
[0088] [Second Embodiment] In this embodiment, an image decoding apparatus that decodes the bitstream generated by the image encoding apparatus according to the first embodiment will be described. Note that, since requirements common to the first embodiment, such as the configuration of the bitstream, are as described in the first embodiment, the description thereof is omitted.
[0089] A functional configuration example of the image decoding apparatus according to the present embodiment will be described with reference to the block diagram of FIG. 2. The separation decoding unit 202 acquires the bitstream generated by the image encoding apparatus according to the first embodiment. The method of acquiring the bitstream is not limited to a specific acquisition method. For example, the bitstream may be acquired directly or indirectly from the image encoding apparatus via a network such as a LAN or the Internet, or the bitstream stored inside or outside the image decoding apparatus may be acquired. Then, the separation decoding unit 202 separates the information related to the decoding process and the coded data related to the coefficients from the acquired bitstream and sends them to the decoding unit 203. Further, the separation decoding unit 202 decodes the coded data of the header of the bitstream. In the present embodiment, the header information related to the division of the image such as the size of the tile, block, slice, and basic block is decoded to generate division information, and the generated division information is output to the image reproduction unit 205. That is, the separation decoding unit 202 performs an operation reverse to that of the integrated encoding unit 111 in FIG. 1.
[0090] The decoding unit 203 decodes the coded data output from the separation decoding unit 202 and reproduces the quantization coefficients and prediction information. The inverse quantization / inverse transformation unit 204 performs inverse quantization on the quantization coefficients to generate transformation coefficients, and performs inverse orthogonal transformation on the generated transformation coefficients to reproduce the prediction error.
[0091] The frame memory 206 is a memory for storing the image data of the reproduced picture. The image reproduction unit 205 appropriately refers to the frame memory 206 based on the input prediction information to generate a predicted image. Then, the image reproduction unit 205 generates a reproduced image from the generated predicted image and the prediction error reproduced by the inverse quantization / inverse transformation unit 204. Then, the image reproduction unit 205 specifies the positions of the tile, block, and slice in the input image of the reproduced image based on the division information input from the separation decoding unit 202 and outputs them.
[0092] The in-loop filter unit 207 performs in-loop filter processing such as deblocking filtering on the reproduced image, similar to the in-loop filter unit 109 described above, and outputs the image on which the in-loop filter processing has been performed. The control unit 299 controls the operation of the entire image decoding apparatus and controls the operation of each functional unit of the image decoding apparatus described above.
[0093] Next, the decoding process of the bitstream by the image decoding apparatus having the configuration shown in FIG. 2 will be described. Hereinafter, it will be described assuming that the bitstream is input to the image decoding apparatus in units of frames, but it may be possible to input the bitstream of a still image for one frame to the image decoding apparatus. Further, in the present embodiment, for the sake of simplicity of explanation, only the intra prediction decoding process will be described, but the present invention is not limited thereto and is also applicable to the inter prediction decoding process.
[0094] The separation decoding unit 202 separates information related to the decoding process and coded data related to coefficients from the input bitstream and sends it to the decoding unit 203. Further, the separation decoding unit 202 decodes the coded data of the header of the bitstream. More specifically, the separation decoding unit 202 decodes the basic block data division information, tile data division information, block data division information, slice data division information 0, basic block row data synchronization information, basic block row data position information, etc. in FIG. 6 to generate division information. Then, the separation decoding unit 202 outputs the generated division information to the image reproduction unit 205. Further, the separation decoding unit 202 reproduces the coded data of the basic block unit of the picture data and outputs it to the decoding unit 203.
[0095] The decoding unit 203 decodes the coded data output from the separation decoding unit 202 and reproduces the quantization coefficient and prediction information. The reproduced quantization coefficient is output to the inverse quantization / inverse transformation unit 204, and the reproduced prediction information is output to the image reproduction unit 205.
[0096] The inverse quantization and inverse transformation unit 204 performs inverse quantization on the input quantization coefficients to generate transformation coefficients, and performs inverse orthogonal transformation on the generated transformation coefficients to reproduce the prediction error. The reproduced prediction error is output to the image reproduction unit 205.
[0097] The image reproduction unit 205 generates a predicted image by appropriately referring to the frame memory 206 based on the prediction information input from the separation and decoding unit 202. Then, the image reproduction unit 205 generates a reproduced image from the generated predicted image and the prediction error reproduced by the inverse quantization and inverse transformation unit 204. Then, for the reproduced image, the image reproduction unit 205 specifies, based on the division information input from the separation and decoding unit 202, the shape and position in the input image of tiles, bricks, slices, etc. as shown in FIG. 7, for example, and outputs (stores) it in the frame memory 206. The image stored in the frame memory 206 is used for reference during prediction.
[0098] The in-loop filter unit 207 performs in-loop filter processing such as deblocking filter on the reproduced image read from the frame memory 206, and outputs (stores) the image subjected to the in-loop filter processing to the frame memory 206.
[0099] The control unit 299 outputs the reproduced image stored in the frame memory 206. The output destination of the reproduced image is not limited to a specific output destination. For example, the control unit 299 may output the reproduced image to a display device included in the image decoding apparatus and display the reproduced image on the display device. Also, for example, the control unit 299 may transmit the reproduced image to an external device via a network such as a LAN or the Internet.
[0100] Next, the decoding process of the bit stream by the image decoding apparatus according to the present embodiment (the decoding process of the bit stream having the configuration of FIG. 6) will be described with reference to the flowchart of FIG. 4.
[0101] In step S401, the separation and decoding unit 202 separates information related to decoding processing and coded data related to coefficients from the input bit stream and sends them to the decoding unit 203. Also, the separation and decoding unit 202 decodes the coded data of the header of the bit stream. More specifically, the separation and decoding unit 202 decodes basic block data division information, tile data division information, block data division information, slice data division information, basic block row data synchronization information, basic block row data position information, etc. in FIG. 6 to generate division information. Then, the separation and decoding unit 202 outputs the generated division information to the image reproduction unit 205. Also, the separation and decoding unit 202 reproduces the coded data of the picture data in units of basic blocks and outputs it to the decoding unit 203.
[0102] In this embodiment, the division of the input image, which is the source of the bit stream coding, is the division shown in FIG. 8. Information related to the input image, which is the source of the bit stream coding, and its division can be derived from the division information.
[0103] From pic_width_in_luma_samples included in the image size information, it can be specified that the horizontal size (width) of the input image is 1152 pixels. Also, from pic_height_in_luma_samples included in the image size information, it can be specified that the vertical size (height) of the input image is 1152 pixels.
[0104] Also, since log2_ctu_size_minus2 = 4 in the basic block data division information, the size of the basic block can be derived as 64×64 pixels from 1<<log2_ctu_size_minus2+2.
[0105] Also, since single_tile_in_pic_flag = 0 in the tile data division information, it can be specified that the input image is divided into a plurality of tiles. And since uniform_tile_spacing_flag = 1, it can be specified that each tile has the same size (excluding the edges).
[0106] Also, since tile_cols_width_minus1 = 5 and tile_rows_height_minus1 = 5, it can be determined that each tile is composed of 6×6 basic blocks. That is, it can be determined that each tile is composed of 384×384 pixels. Since the input image is 1152×1152 pixels, it can be seen that the input image is divided into 9 tiles, 3 in the horizontal direction and 3 in the vertical direction, and encoded.
[0107] Also, since brick_splitting_present_flag of the brick data division information is 1, it can be determined that at least one tile in the input image is divided into a plurality of bricks.
[0108] Also, brick_split_flag[1], brick_split_flag[4], brick_split_flag[5], brick_split_flag[6], brick_split_flag[8] are 0. Thereby, it can be determined that the tiles corresponding to tile ID = 1, 4, 5, 6, 8 are not divided into bricks. In this embodiment, since the number of basic block rows of all tiles is 6, it can be seen that the number of basic block rows of the bricks of the tiles corresponding to tile ID = 1, 4, 5, 6, 8 is 6.
[0109] On the one hand, brick_split_flag[0], brick_split_flag[2], brick_split_flag[3], and brick_split_flag[7] are all 1. Thus, it can be specified that the tiles corresponding to tile IDs = 0, 2, 3, and 7 are split into bricks. Also, uniform_brick_spacing_flag[0], uniform_brick_spacing_flag[3], and uniform_brick_spacing_flag[7] are all 1. Thus, it can be specified that the tiles corresponding to tile IDs = 0, 3, and 7 are all split into bricks of the same size.
[0110] Also, brick_height_minus1[0] and brick_height_minus1[7] are both 2. Then, it can be specified that for the tile corresponding to tile ID = 0 and the tile corresponding to tile ID = 7, the number of basic blocks in the vertical direction of the bricks within the tile is 3. Also, for the tile corresponding to tile ID = 0 and the tile corresponding to tile ID = 7, it can be specified that the number of bricks within the tile is 2 (= the number of basic block rows of the tile (6) / the number of basic blocks in the vertical direction of the bricks within the tile (3)).
[0111] Also, brick_height_minus1[3] is 1. Then, it can be specified that the number of basic blocks in the vertical direction of the bricks within the tile corresponding to tile ID = 3 is 2. Also, it can be specified that the number of bricks within the tile corresponding to tile ID = 3 is 3 (= the number of basic block rows of the tile (6) / the number of basic blocks in the vertical direction of the bricks within the tile (2)).
[0112] For the tile corresponding to Tile ID = 2, since num_brick_rows_minus1[2] = 1, it can be determined that the tile is composed of two bricks. Also, uniform_brick_spacing_flag[2] = 0. Thus, it can be determined that there are bricks of different sizes in the tile corresponding to Tile ID = 2. And brick_row_height_minus1[2][0] = 1, brick_row_height_minus1[2][1] = 3, and the number of basic blocks in the vertical direction of all tiles is 6. Thus, it can be determined that the tile corresponding to Tile ID = 2 is composed of a brick with 2 basic blocks in the vertical direction and a brick with 4 basic blocks in the vertical direction, in order from the top. Note that brick_row_height_minus1[2][1] = 3 may not be encoded. This is because when the number of bricks in a tile is two, it is possible to obtain the height of the second brick from the height of the tile and the height of the first brick in the tile (brick_row_height_minus1[2][0] = 1).
[0113] Also, since single_brick_per_slice_flag = 0 in the slice data division information 0, it can be determined that at least one slice is composed of multiple bricks. In this embodiment, when uniform_brick_spacing_flag[i] = 0, num_brick_rows_minus1[i], which indicates (the number of bricks constituting the i-th tile - 1), is included in the brick data division information. However, it is not limited to this.
[0114] For example, assume that when brick_split_flag[i]=1, the number of bricks constituting the i-th tile is 2 or more. Then, num_brick_rows_minus2[i], which indicates (the number of bricks constituting the tile - 2), may be decoded instead of num_brick_rows_minus1[i]. By doing so, it is possible to decode a bit stream with the number of bits of the syntax indicating the number of bricks constituting the tile reduced.
[0115] Next, obtain the coordinates of the upper left and lower right boundaries of each brick. The coordinates are based on the upper left of the input image as the origin and are indicated by the horizontal and vertical positions of the basic block. For example, the coordinates of the upper left boundary of the third basic block from the left and the second basic block from the top are (3, 2), and the coordinates of the lower right boundary are (4, 3).
[0116] The coordinates of the upper left boundary of the brick with BID = 0 within the tile corresponding to tile ID = 0 are (0, 0). Since the number of basic block rows of the brick with BID = 0 is 3 and the number of basic blocks in the horizontal direction of all tiles is 6, the coordinates of the lower right boundary are (3, 3).
[0117] The coordinates of the upper left boundary of the brick with BID = 1 within the tile corresponding to tile ID = 0 are (0, 3). Since the number of basic block rows of the brick with BID = 1 is 3 and the number of basic blocks in the horizontal direction of all tiles is 6, the coordinates of the lower right boundary are (6, 6).
[0118] The coordinates of the upper left boundary of the tile (brick with BID = 2) corresponding to tile ID = 1 are (6, 0). Since the number of basic block rows of the brick with BID = 2 is 6 and the number of basic blocks in the horizontal direction of all tiles is 6, the coordinates of the lower right boundary are (12, 6).
[0119] The coordinates of the upper left boundary of the brick with BID = 3 within the tile corresponding to tile ID = 2 are (12, 0). Since the number of basic block rows of the brick with BID = 3 is 2, the coordinates of the lower right boundary are (18, 2).
[0120] The coordinates of the upper left boundary of the block with BID = 4 in the tile corresponding to tile ID = 2 are (12, 2). Since the number of basic block rows of the block with BID = 4 is 4, the coordinates of the lower right boundary are (18, 6).
[0121] The coordinates of the upper left boundary of the block with BID = 5 in the tile corresponding to tile ID = 3 are (0, 6). Since the number of basic block rows of the block with BID = 5 is 2, the coordinates of the lower right boundary are (6, 8).
[0122] The coordinates of the upper left boundary of the block with BID = 6 in the tile corresponding to tile ID = 3 are (0, 8). Since the number of basic block rows of the block with BID = 6 is 2, the coordinates of the lower right boundary are (6, 10).
[0123] The coordinates of the upper left boundary of the block with BID = 7 in the tile corresponding to tile ID = 3 are (0, 10). Since the number of basic block rows of the block with BID = 7 is 2, the coordinates of the lower right boundary are (6, 12).
[0124] The coordinates of the upper left boundary of the tile (block with BID = 8) corresponding to tile ID = 4 are (6, 6). Since the number of basic block rows of the block with BID = 8 is 6 and the number of basic blocks in the horizontal direction of all tiles is 6, the coordinates of the lower right boundary are (12, 12).
[0125] The coordinates of the upper left boundary of the tile (block with BID = 9) corresponding to tile ID = 5 are (12, 6). Since the number of basic block rows of the block with BID = 9 is 6 and the number of basic blocks in the horizontal direction of all tiles is 6, the coordinates of the lower right boundary are (18, 12).
[0126] The coordinates of the upper left boundary of the tile (block with BID = 10) corresponding to tile ID = 6 are (0, 12). Since the number of basic block rows of the block with BID = 10 is 6 and the number of basic blocks in the horizontal direction of all tiles is 6, the coordinates of the lower right boundary are (6, 18).
[0127] The coordinates of the upper left boundary of the brick with BID = 11 in the tile corresponding to tile ID = 7 are (6, 12). Since the number of basic block rows of the brick with BID = 11 is 3 and the number of basic blocks in the horizontal direction of all tiles is 6, the coordinates of the lower right boundary are (12, 15).
[0128] The coordinates of the upper left boundary of the brick with BID = 12 in the tile corresponding to tile ID = 7 are (6, 15). Since the number of basic block rows of the brick with BID = 12 is 3 and the number of basic blocks in the horizontal direction of all tiles is 6, the coordinates of the lower right boundary are (12, 18).
[0129] The coordinates of the upper left boundary of the tile (brick with BID = 13) corresponding to tile ID = 8 are (12, 12). Since the number of basic block rows of the brick with BID = 13 is 6 and the number of basic blocks in the horizontal direction of all tiles is 6, the coordinates of the lower right boundary are (18, 18).
[0130] Next, identify the bricks included in each slice. Since num_slices_in_pic_minus1 = 4, it can be determined that the number of slices in the input image is 5. Also, the corresponding brick can be identified from the slice_address of the slice to be processed. That is, when the slice_address is N, it can be known that the slice to be processed is slice N.
[0131] In the case of slice 0, bottom_right_brick_idx_delta[0] is 2. Thus, it can be determined that the bricks included in slice 0 are the bricks included in the rectangular region surrounded by the coordinates of the upper left boundary of the brick with BID = 0 and the coordinates of the lower right boundary of the brick with BID = 2. Since the coordinates of the upper left boundary of the brick with BID = 0 are (0, 0) and the lower right boundary of the brick with BID = 2 is (12, 6), it can be determined that the bricks included in slice 0 are the bricks with BID = 0 to 2.
[0132] In the case of slice 1, since top_left_brick_idx[1] is 3 and bottom_right_brick_idx_delta[1] is 0, it can be determined that the bricks included in slice 1 are the bricks with BID = 3.
[0133] In the case of slice 2, since top_left_brick_idx[2] is 4 and bottom_right_brick_idx_delta[2] is 0, it can be determined that the bricks included in slice 2 are the bricks with BID = 4.
[0134] In the case of slice 3, top_left_brick_idx[3] is 5 and bottom_right_brick_idx_delta[3] is 7. Since the coordinates of the upper left boundary of the brick with BID = 5 are (0, 6) and the coordinates of the lower right boundary of the brick with BID = 12 are (12, 18), it can be determined that the bricks included in slice 3 are those in the area with the upper left coordinates (0, 6) and the lower right coordinates (12, 18). As a result, it can be determined that the bricks corresponding to BID = 5 to 8 and 10 to 12 are the bricks included in slice 3. The coordinates of the lower right boundary of the brick corresponding to BID = 9 are (18, 12), and the coordinates of the lower right boundary of the brick corresponding to BID = 13 are (18, 18), both of which are outside the range of slice 3, so it is determined that they do not belong to slice 3.
[0135] In the case of slice 4, top_left_brick_idx[4] is 9 and bottom_right_brick_idx_delta[4] is 4. Since the coordinates of the upper left boundary of the brick with BID = 9 are (12, 6) and the coordinates of the lower right boundary of the brick with BID = 13 are (18, 18), it can be determined that the bricks included in the region with the upper left coordinates (12, 6) and the lower right coordinates (18, 18) are the bricks included in slice 4. As a result, it can be determined that the bricks corresponding to BID = 9 and 13 are the bricks included in slice 4. Here, the coordinates of the upper left boundary of the brick corresponding to BID = 10 are (0, 12), the coordinates of the upper left boundary of the brick corresponding to BID = 11 are (6, 12), and the coordinates of the upper left boundary of the brick corresponding to BID = 12 are (6, 15). Therefore, since all of them are outside the range of slice 4, it is determined that they do not belong to slice 4.
[0136] In this embodiment, the bricks included in slice 4, which is the last slice of the input image, are specified from top_left_brick_idx[4] and bottom_right_brick_idx_delta[4], but it is not limited to this. It has already been derived that the bricks included in slices 0 to 3 are the bricks with BID = 0 to 2, BID = 3, BID = 4, BID = 5 to 8, and 10 to 12 respectively, and it has already been derived that the input image is composed of 14 bricks with BID = 0 to 13. Therefore, it is possible to specify that the bricks included in the last slice are the remaining bricks with BID = 9 and 13. Therefore, even if top_left_brick_idx[4] and bottom_right_brick_idx_delta[4] are not encoded, it is possible to specify the bricks included in the last slice. In this way, it is possible to decode a bit stream with a reduced amount of bits in the header part.
[0137] Also, entropy_coding_sync_enabled_flag = 1 in the basic block row data synchronization information. From this, it can be seen that entry_point_offset_minus1[j - 1] indicating (the size of the coded data of the (j - 1)-th basic block row in the slice - 1) is coded in the bitstream. The number thereof is the number of basic block rows of the slice to be processed - 1.
[0138] In this embodiment, as described above, since the blocks belonging to each slice can be specified, entry_point_offset_minus1[j] is coded by the number obtained by subtracting 1 from the total value of the number of basic block rows of the blocks belonging to the slice.
[0139] In the case of slice 0, the total value (3 + 3 + 6 = 12) of the number of basic block rows of the block corresponding to BID = 0 (3) + the number of basic block rows of the block corresponding to BID = 1 (3) + the number of basic block rows of the block corresponding to BID = 2 (6) - 1 = 11. Therefore, for slice 0, 11 entry_point_offset_minus1[] are coded. In this case, the range of j is 0 to 10.
[0140] In the case of slice 1, the number of basic block rows of the block corresponding to BID = 3 (2) - 1 = 1. Therefore, for slice 1, 1 entry_point_offset_minus1[] is coded. In this case, the range of j is only 0.
[0141] In the case of slice 2, the number of basic block rows of the block corresponding to BID = 4 (4) - 1 = 3. Therefore, for slice 3, 3 entry_point_offset_minus1[] are coded. In this case, the range of j is 0 to 2.
[0142] In the case of slice 3, the total value of the number of basic block rows of the blocks corresponding to each of BID = 5 to 8 and 10 to 12 (2 + 2 + 2 + 6 + 6 + 3 + 3 = 24) - 1 = 23. Therefore, for slice 3, 23 entry_point_offset_minus1[] are encoded. In this case, the range of j is 0 to 22.
[0143] In the case of slice 4, the total value of the number of basic block rows of the block corresponding to BID = 9 (6) + the number of basic block rows of the block corresponding to BID = 13 (6) (6 + 6 = 12) - 1 = 11. Therefore, for slice 4, 11 entry_point_offset_minus1[] are encoded. In this case, the range of j is 0 to 10.
[0144] As described above, even without encoding num_entry_point_offset as in the conventional method, the number of entry_point_offset_minus1 can be derived from other syntaxes. And since the start position of the data of each basic block row is known, the decoding process can be performed in parallel for each basic block row. The division information derived from the separation decoder 202 is sent to the image reproduction unit 205 and used to specify the position of the processing target in the input image in step S404.
[0145] In step S402, the decoder 203 decodes the coded data separated by the separation decoder 202 and reproduces the quantization coefficients and prediction information. In step S403, the inverse quantization / inverse transform unit 204 performs inverse quantization on the input quantization coefficients to generate transform coefficients, and performs inverse orthogonal transform on the generated transform coefficients to reproduce the prediction error.
[0146] In step S404, the image playback unit 205 appropriately refers to the frame memory 206 based on the prediction information input from the decoding unit 203 to generate a predicted image. Then, the image playback unit 205 generates a reproduced image from the generated predicted image and the prediction error reproduced by the inverse quantization / inverse transformation unit 204. Then, for the reproduced image, the image playback unit 205 specifies the positions of the tiles, blocks, and slices in the input image based on the division information input from the separation decoding unit 202, synthesizes them at those positions, and outputs (stores) them in the frame memory 206.
[0147] In step S405, the control unit 299 determines whether all the basic blocks of the input image have been decoded. As a result of this determination, if all the basic blocks of the input image have been decoded, the process proceeds to step S406. On the other hand, if there are still basic blocks in the input image that have not been decoded, the process proceeds to step S402, and decoding processing is performed on the basic blocks that have not yet been decoded.
[0148] In step S406, the in-loop filter unit 207 performs in-loop filter processing on the reproduced image read from the frame memory 206, and outputs (stores) the image subjected to the in-loop filter processing in the frame memory 206.
[0149] Thus, according to the present embodiment, the input image can be decoded from the "bit stream that does not include information indicating how many pieces of information indicating the head positions of the basic block rows included in the block are encoded" generated by the image encoding apparatus according to the first embodiment.
[0150] Note that the image encoding apparatus according to the first embodiment and the image decoding apparatus according to the second embodiment may be separate apparatuses. Also, the image encoding apparatus according to the first embodiment and the image decoding apparatus according to the second embodiment may be integrated into one apparatus.
[0151] [Third Embodiment] Each functional unit shown in FIG. 1 or FIG. 2 may be implemented in hardware, but a part of it may also be implemented in software. In the latter case, each functional unit except for the frame memory 108 and the frame memory 206 may be implemented in software (computer program). A computer device capable of executing such a computer program is applicable to the above-described image encoding device and image decoding device.
[0152] An example of the hardware configuration of a computer device applicable to the above-described image encoding device and image decoding device will be described with reference to the block diagram of FIG. 5. Note that the hardware configuration shown in FIG. 5 is merely an example of the hardware configuration of a computer device applicable to the above-described image encoding device and image decoding device, and can be appropriately changed / modified.
[0153] The CPU 501 executes various processes using computer programs and data stored in the RAM 502 and the ROM 503. Thereby, the CPU 501 controls the operation of the entire computer device, and executes or controls each process described as being performed by the above-described image encoding device and image decoding device. That is, the CPU 501 can function as each functional unit (excluding the frame memory 108 and the frame memory 206) shown in FIG. 1 and FIG. 2.
[0154] The RAM 502 has an area for storing computer programs and data loaded from the ROM 503 and the external storage device 506, and an area for storing data received from the outside via the I / F 507. The RAM 502 also has a work area used when the CPU 501 executes various processes. In this way, the RAM 502 can appropriately provide various areas. The ROM 503 stores setting data of the computer device, a startup program, and the like.
[0155] The operation unit 504 is a user interface such as a keyboard, a mouse, or a touch panel screen, and various instructions can be input to the CPU 501 by a user's operation.
[0156] The display unit 505 is composed of a liquid crystal screen, a touch panel screen, etc., and can display the processing results by the CPU 501 in the form of images, characters, etc. Note that the display unit 505 may be a device such as a projector that projects images and characters.
[0157] The external storage device 506 is a large-capacity information storage device such as a hard disk drive device. In the external storage device 506, an OS (operating system), and computer programs and data for causing the CPU 501 to execute or control each of the above-described processes performed by the above-described image encoding device and image decoding device are stored.
[0158] The computer programs stored in the external storage device 506 include computer programs for causing the CPU 501 to execute or control the functions of each functional unit except for the frame memories 108 and 206 in FIGS. 1 and 2. Further, the data stored in the external storage device 506 includes those described as known information in the above description and various information related to encoding and decoding.
[0159] The computer programs and data stored in the external storage device 506 are appropriately loaded into the RAM 502 according to the control by the CPU 501 and become the processing targets by the CPU 501.
[0160] The frame memory 108 included in the image encoding device in FIG. 1 and the frame memory 206 included in the image encoding device in FIG. 2 can be implemented using memory devices such as the above-described RAM 502 and external storage device 506.
[0161] I / F 507 is an interface for performing data communication with an external device. For example, when a computer device is applied to an image encoding device, the image encoding device can output the generated bitstream externally via I / F 507. Also, when a computer device is applied to an image decoding device, the image decoding device can receive the bitstream via I / F 507. Further, the image decoding device can transmit externally via I / F 507 the result of decoding the bitstream. The CPU 501, RAM 502, ROM 503, operation unit 504, display unit 505, external storage device 506, and I / F 507 are all connected to the bus 508.
[0162] Note that the specific numerical values used in the above description are for the purpose of specific explanation and are not intended to limit each of the above embodiments to these numerical values. Also, it is possible to appropriately combine some or all of the above-described embodiments. Further, it is possible to selectively use some or all of the above-described embodiments.
[0163] (Other Embodiments) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. Also, it can be realized by a circuit (for example, ASIC) that realizes one or more functions.
[0164] The invention is not limited to the above embodiments, and various changes and modifications are possible without departing from the spirit and scope of the invention. Therefore, claims are attached to disclose the scope of the invention.
Explanation of Reference Numerals
[0165] 102: Image segmentation unit 103: Block segmentation unit 104: Prediction unit 105: Transformation / quantization unit 106: Inverse quantization / inverse transformation unit 107: Image reproduction unit 108: Frame memory 109: In-loop filter unit 110: Encoding unit 111: Integrated encoding unit
Claims
1. An image decoding apparatus for decoding an image obtained by encoding an image including a rectangular area including one or more block rows composed of a plurality of blocks from a bit stream, a decoding means for decoding from the bit stream first information indicating an integer value n equal to the result of subtracting 1 from the number of slices included in the image, a first flag related to enabling parallel processing, and second information used for specifying a first rectangular area to be processed among a plurality of rectangular areas included in a target slice which is a slice corresponding to the i-th (i is a certain integer value) in the image, third information used for specifying a last rectangular area to be processed among the plurality of rectangular areas, and fourth information corresponding to the number of blocks in the vertical direction of the rectangular area in the image; a specifying means for specifying, when the value of the first flag is 1 and a second flag which is a flag decoded from a picture parameter set of the bit stream and related to a slice mode indicates that a mode in which a slice is rectangular is used, and in a state where the target slice included in the image includes a plurality of rectangular areas at least in the horizontal direction or the vertical direction, the number of syntax elements used for specifying the head position of coded data of a block row for the target slice based on at least the second information, the third information, and the fourth information; comprising; when the target slice in the image is the slice corresponding to the n-th slice, the third information is not decoded from the bit stream; the number of syntax elements specified by the specifying means is included in a slice header of the bit stream; intra prediction is available for blocks in the image An image decoding apparatus characterized by the above.
2. The decoding means is capable of decoding each block row in the rectangular area in parallel. The image decoding apparatus according to Claim 1.
3. The first flag is an entropy_coding_sync_enabled_flag. The image decoding apparatus according to Claim 1.
4. The second flag is a rect_slice_flag. The image decoding apparatus according to Claim 1.
5. The image decoding apparatus according to claim 1, wherein information on the number of the syntax elements is not signaled in the bit stream.
6. The image decoding apparatus according to claim 1, wherein the second information is decoded from the picture parameter set of the bit stream.
7. The image decoding apparatus according to claim 1, wherein the third information is decoded from the picture parameter set of the bit stream.
8. The image decoding apparatus according to claim 1, wherein the fourth information is decoded from the picture parameter set of the bit stream.
9. In the image decoding apparatus according to claim 1, among the plurality of rectangular regions included in the target slice, the rectangular region processed first is the upper left corner rectangular region among the plurality of rectangular regions, and the rectangular region processed last among the plurality of rectangular regions included in the target slice is the lower right corner rectangular region among the plurality of rectangular regions.
10. The image decoding apparatus according to claim 1, wherein the size of each of the plurality of blocks forming the block row is determined from fifth information decoded from a sequence parameter set in the bit stream.
11. The image decoding apparatus according to claim 10, wherein the size of each of the plurality of blocks is derived by arithmetically left-shifting 1 by a value which is a result of adding a predetermined value to the fifth information.
12. The image decoding apparatus according to claim 1, wherein the image includes a plurality of slices from the 0th slice to the nth slice.
13. The image decoding apparatus according to claim 12, wherein the 0th slice is the upper left corner slice of the image, and the nth slice is the lower right corner slice of the image.
14. The image decoding apparatus according to claim 1, wherein each of the plurality of blocks forming the block row corresponds to a CTU (Coding Tree Unit) and is dividable into a plurality of small blocks.
15. An image decoding method for decoding an image from a bit stream obtained by encoding an image including a rectangular region including one or more block rows composed of a plurality of blocks, A decoding step of decoding from the bitstream first information indicating an integer value n equal to the result of subtracting 1 from the number of slices included in the image, a first flag related to enabling parallel processing, second information used to identify the first rectangle region to be processed among a plurality of rectangle regions included in a target slice that is the slice corresponding to the i-th (i is a certain integer value) in the image, third information used to identify the last rectangle region to be processed among the plurality of rectangle regions, and fourth information corresponding to the number of blocks in the vertical direction of the rectangle region in the image; A specifying step of specifying, in a state where the value of the first flag is 1, the first flag is a flag decoded from a picture parameter set of the bitstream and related to the slice mode, the second flag indicating that a mode in which the slice is rectangular is used, and the target slice included in the image includes at least a plurality of rectangle regions in the horizontal or vertical direction, the number of syntax elements used to specify the start position of the coded data of the block row for the target slice based on at least the second information, the third information, and the fourth information; having; when the target slice in the image is the slice corresponding to the n-th, the third information is not decoded from the bitstream; the number of syntax elements specified by the specifying step is included in the slice header of the bitstream; intra prediction is available for blocks in the image An image decoding method characterized by this.
16. An image encoding apparatus for encoding an image including a rectangular region including one or more block rows composed of a plurality of blocks, Encoding means for encoding into a bitstream first information indicating an integer value n equal to the result of subtracting 1 from the number of slices included in the image, a first flag related to enabling parallel processing, second information used to identify the first rectangle region to be processed among a plurality of rectangle regions included in a target slice that is the slice corresponding to the i-th (i is a certain integer value) in the image, third information used to identify the last rectangle region to be processed among the plurality of rectangle regions, and fourth information corresponding to the number of blocks in the vertical direction of the rectangle region in the image; The value of the first flag is 1, and a second flag, which is a flag encoded in the picture parameter set of the bit stream and is related to the mode of a slice, indicates that a mode in which the slice is rectangular is used. Also, in a state where the target slice included in the image includes a plurality of rectangular regions at least in the horizontal direction or the vertical direction, based on at least the second information, the third information, and the fourth information, specific means for specifying the number of syntax elements used to specify the start position of the coded data of a block row for the target slice comprising when the target slice in the image is the slice corresponding to the nth slice, the third information is not encoded in the bit stream, the number of syntax elements specified by the specific means is included in the slice header of the bit stream, intra prediction is available for a block in the image characterized by an image encoding apparatus.
17. An image encoding method for encoding an image including a rectangular region including one or more block rows composed of a plurality of blocks, a coding step of coding in a bit stream first information indicating an integer value n equal to the result of subtracting 1 from the number of slices included in the image, a first flag related to enabling parallel processing, second information used to specify the first rectangular region to be processed among a plurality of rectangular regions included in a target slice which is the slice corresponding to the i-th (i is a certain integer value) in the image, third information used to specify the last rectangular region to be processed among the plurality of rectangular regions, and fourth information corresponding to the number of blocks in the vertical direction of the rectangular region in the image; the value of the first flag is 1, and a second flag, which is a flag encoded in the picture parameter set of the bit stream and is related to the mode of a slice, indicates that a mode in which the slice is rectangular is used. Also, in a state where the target slice included in the image includes a plurality of rectangular regions at least in the horizontal direction or the vertical direction, based on at least the second information, the third information, and the fourth information, a specific step of specifying the number of syntax elements used to specify the start position of the coded data of a block row for the target slice having When the target slice in the image is the slice corresponding to the nth slice, the third information is not encoded in the bitstream. The number of syntax elements specified by the specifying step is included in the slice header of the bitstream. Intra prediction is available for blocks in the image. An image encoding method characterized by the above.
18. A computer program for causing a computer to execute the image decoding method according to claim 15.
19. A computer program for causing a computer to execute the image encoding method according to claim 17.
Citation Information
Patent Citations
Image encoder, image encoding method and program, image decoder, and image decoding method and program
JP2014011638A